Clustering Method, Device and Server for Data Objects
By adopting clustering methods in data processing scenarios, using feature data and multi-party security calculations, the problem of how data parties protect data privacy during clustering is solved, and a safe and efficient user clustering is achieved.
Patent Information
- Application Number
- CN202110817153.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-20
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-07-20
AI Technical Summary
In data processing scenarios, two different data parties need to cooperate to divide users into corresponding type groups, and in this process, it is necessary to avoid leaking the data they hold to the other party to protect data privacy.
Through a clustering method of data objects, using preset protocol rules, they interact with the other party using the feature data of the data objects they hold, calculate the fragment of the feature distance, and determine the matching center object of the data object through multi-party security calculations, and finally divide the data object into the corresponding group.
It realizes that without leaking our own data to the other party, safely cooperates to complete user clustering, protects the data privacy of both parties, and reduces the risk of data leakage.
Smart Images

Figure CN113657451B_ABST
Abstract
Description
Technical Field
[0001] This specification belongs to the field of Internet technologies, and particularly relates to a method, apparatus, and server for clustering data objects. Background Art
[0002] In some data processing scenarios, sometimes two different data parties may hold different types of feature data of the same batch of users. Currently, the two parties need to cooperate to classify this batch of users into corresponding type groups by using the data they hold respectively; and it is also required to avoid disclosing the data held by themselves to the other party during the classification process to protect the data privacy of the participating parties.
[0003] Therefore, there is an urgent need for a method that can securely cooperate to complete user clustering without disclosing the data held by oneself to the other party. Summary of the Invention
[0004] This specification provides a method, apparatus, and server for clustering data objects, which can securely cooperate to complete the clustering of data objects without disclosing the data held by oneself to the other party.
[0005] The method, apparatus, and server for clustering data objects provided in this specification are implemented as follows:
[0006] A method for clustering data objects includes: according to a preset protocol rule, using the first type of feature data of the data object held and the first type of feature data of the first central object, performing a first interaction with a second data party to obtain a first shard of the feature distance between the data object and the first central object; wherein, the first central object is the class group center determined in the previous clustering; the second data party holds the second type of feature data of the data object and the second type of feature data of the first central object; the second data party obtains a second shard of the feature distance between the data object and the first central object by performing a first interaction with the first data party; according to the preset protocol rule, using the first shard of the feature distance between the data object and the first central object, performing a second interaction with the second data party to determine a matching central object of the data object; wherein, the matching central object is the first central object with the smallest feature distance from the data object; and classifying the data object into the class group corresponding to the matching central object.
[0007] A clustering device for data objects, comprising: a first interaction module, configured to perform a first interaction with a second data party according to a preset protocol rule, using first-class feature data of the data objects it holds and first-class feature data of a first central object, so as to obtain a first fragment of the feature distance between the data objects and the first central object; wherein, the first central object is the class group center determined in the previous clustering; the second data party holds second-class feature data of the data objects and second-class feature data of the first central object; the second data party obtains a second fragment of the feature distance between the data objects and the first central object by performing a first interaction with a first data party; a second interaction module, configured to perform a second interaction with the second data party according to a preset protocol rule, using the first fragment of the feature distance between the data objects and the first central object, so as to determine a matching central object of the data objects; wherein, the matching central object is the first central object with the smallest feature distance from the data objects; a classification module, configured to classify the data objects into the class group corresponding to the matching central object.
[0008] A server, comprising a processor and a memory for storing processor-executable instructions, wherein when the processor executes the instructions, relevant steps of the clustering method for data objects are implemented.
[0009] A computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed, relevant steps of the clustering method for data objects are implemented.
[0010] A clustering method, device and server for data objects provided in this specification. When a first data party holding first-class feature data of data objects and a second data party holding second-class feature data of data objects cooperate to cluster the data objects, they can first perform a first interaction using the data they hold respectively according to a preset protocol rule, so as to obtain a fragment of the feature distance between the data objects and a first central object respectively; then, according to the preset protocol rule, they can perform a second interaction using the fragment of the feature distance they hold respectively, so as to determine the matching central object of each data object from the first data objects; and then classify the data objects into the class group corresponding to the matching central object to complete one clustering of the data objects. Thereby, it can enable the first data party and the second data party to safely and efficiently cooperate to complete the clustering of data objects without disclosing the data they hold to each other, protect the data privacy of both parties, and reduce the risk of the data they hold being leaked during the clustering process. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] To more clearly illustrate the embodiments of this specification, the accompanying drawings required for use in the embodiments will be briefly introduced below. The accompanying drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0012] Figure 1 It is a schematic diagram of an embodiment of the structural composition of a system applying the clustering method of data objects provided by the embodiments of this specification;
[0013] Figure 2 It is a schematic flowchart of the clustering method of data objects provided by an embodiment of this specification;
[0014] Figure 3 It is a schematic diagram of an embodiment applying the clustering method of data objects provided by the embodiments of this specification in a scenario example;
[0015] Figure 4 It is a schematic diagram of the structural composition of a server provided by an embodiment of this specification;
[0016] Figure 5 It is a schematic diagram of the structural composition of a clustering device for data objects provided by an embodiment of this specification. Detailed implementation manners
[0017] To enable those skilled in the art of this technology to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this specification.
[0018] The embodiments of this specification provide a clustering method for data objects. The clustering method for data objects can be specifically applied to a system including a first server and a second server. Specifically, reference can be made to Figure 1 as shown. The first server and the second server can be connected by wired or wireless means for specific data interaction.
[0019] Among them, the above-mentioned first server can be specifically understood as a server deployed on the side of the first data party (for example, a certain shopping website, etc.). Specifically, the first server can at least hold the first type of feature data owned by the first data party. For example, the shopping-related features of the user object can specifically include: the shopping frequency of the user object, the monthly shopping consumption amount of the user object, the types of goods purchased by the user object, the brand of goods most frequently purchased by the user object, and so on.
[0020] The above-mentioned second server can be specifically understood as a server deployed on the side of the second data party (for example, a certain bank, etc.). Specifically, the second server can at least hold the second type of feature data owned by the second data party. For example, the asset-related features of the user object can specifically include: the monthly income of the user object, the mortgage record of the user object, the financial management income of the user object, and so on.
[0021] In this embodiment, the above-mentioned first server and second server can specifically include a background server applied to the side of the data processing system, which can implement functions such as data transmission and data processing. Specifically, the above-mentioned first server and second server can, for example, be an electronic device with data operation, storage functions, and network interaction functions. Or, the above-mentioned first server and second server can also be software programs running in the electronic device, providing support for data processing, storage, and network interaction. In this embodiment, the number of servers included in the above-mentioned first server and second server is not specifically limited. The above-mentioned first server and second server can specifically be one server, or several servers, or a server cluster formed by several servers.
[0022] It should be noted that the feature data of the user object is vertically distributed. Among them, the first type of feature data held by the first server and the second type of feature data held by the second server are the feature data of the same batch of user objects.
[0023] Moreover, in advance, the first server and the second server also performed alignment processing on the first type of feature data and the second type of feature data according to the user objects of the feature data. Specifically, for example, the first type of feature data ranked at the i-th position held by the first server corresponds to the same user object as the second type of feature data ranked at the i-th position held by the second server.
[0024] In addition, the first type of feature data of each data object can specifically include multiple first features, and the multiple first features in each first type of feature data are arranged in a preset first order (for example, arranged in the order that the shopping frequency of the user object ranks first, the monthly shopping consumption amount of the user object ranks second, the types of goods purchased by the user object ranks third, etc.).
[0025] For the first type of feature data of any user object, it can be specifically represented as the following n-dimensional feature vector: x = (x1, x2, x3,..., x n ). Among them, x1, x2, x3,..., x n represent n first features arranged in a preset first order, and n is an integer greater than or equal to 1.
[0026] Similarly, the second type of feature data of each data object can also specifically include multiple second features, and the multiple second features in each second type of feature data are arranged in a preset second order (for example, arranged in the order that the monthly income of the user object ranks first, the mortgage record of the user object ranks second, the financial management income of the user object ranks third, etc.).
[0027] For the second type of feature data of any user object, it can be specifically represented as the following m-dimensional feature vector: y = (y1, y2, y3,..., y m ). Among them, y1, y2, y3,..., y m represent m second features arranged in a preset second order, and m is an integer greater than or equal to 1.
[0028] Actually, corresponding to this user object, the complete feature data should include the first type of feature data and the second type of feature data, which can be recorded as: z = (x1, x2, x3,..., x n , y1, y2, y3,..., y m ), that is, it contains n + m features.
[0029] Currently, a certain shopping website plans to cooperate with a certain bank on the premise of protecting data privacy, each using the data it holds, and jointly using the shopping-related features and asset-related features of user objects to cluster multiple user objects into multiple different groups, so that the shopping website can subsequently distinguish users in different groups and push products targeted.
[0030] Specifically in implementation, a clustering request for user objects can be initiated by the first server or the second server. Among them, the clustering request can specifically carry the identity identifier of the user object to be clustered. For example, the name of the user object, the account name, the registered email, the user number, and so on.
[0031] First, the first server and the second server can respond to the clustering request and determine, through negotiation or other means, that the number of clusters is M. That is, it is determined to cluster user objects into M different groups. Among them, the specific value of the above-mentioned number of clusters M can be flexibly determined according to the specific application scenario and the purchase potential type division rule of the customer group of this shopping website. In this embodiment, the value of M can be set to 3.
[0032] Next, the first server can randomly generate the feature data of M initial first central objects according to the number of clusters. Among them, the feature data of the initial first central object numbered i can be expressed in the following form: <c i >1 = ( <u1> 1, <u2>1,...,<u n >1,...<u n+m >1). Among them, i is an integer greater than or equal to 1 and less than or equal to M. The feature with subscript 1 of the above n + m symbols "<>" can be a random number randomly generated by the first server according to a random number seed.
[0033] At the same time, the second server randomly generates the feature data of M initial second central objects according to the number of clusters. For example, the feature data of the initial second central object numbered i can be expressed in the following form: <c i >2 = ( <u1> 2, <u2>2,...,<u n >2,...<u n+m >2). Among them, i is an integer greater than or equal to 1 and less than or equal to M. The feature with subscript 2 of the above n + m symbols "<>" can specifically be a random number randomly generated by the second server according to a random number seed.
[0034] In some cases, the first server and the second server can share the feature data of the initial first central object and the feature data of the initial second central object.
[0035] Then, the first server can, according to a preset protocol rule (for example, a protocol rule based on secret sharing and oblivious transfer), use the first type of feature data of the user object it holds, and the feature data of the initial first central object, to perform a fifth interaction with the second data party (using the second type of feature data of the user object it holds, and the feature data of the initial second central object), to calculate the first shard (for example, a share) of the feature distance between the user object and the initial central object. At the same time, the second server can calculate the second shard (for example, another share) of the feature distance between the user object and the initial central object.
[0036] Among them, the first shard and the second shard of the above feature distance can be combined to obtain the complete feature distance from the user object to the initial central object.
[0037] The above initial central object can specifically be obtained by combining the initial first central object and the initial second central object. Specifically, the feature data of the initial central object can be expressed in the following form: c i =( <u1> 1+ <u1> 2, <u2> 1+ <u2> 2,..., <u1> t +<u t >2,...<u n+m >1+<u n+m >2).
[0038] Specifically, taking the calculation of the feature distance between any current user object among multiple user objects and the initial central object numbered i as an example, it can be denoted as: |c i -z| = |(<c i >1+<c i >2)-z|. Among them, |c i -z| represents the feature distance between the current user object and the initial central object numbered i, c i represents the feature data of the initial central object numbered i, <c i >1 represents the feature data of the initial first central object numbered i, <c i >2 represents the feature data of the initial second central object numbered i, and z represents the feature data of the current user object (including the first type of feature data and the second type of feature data).
[0039] The first server can first calculate the first intermediate data and the square of the first intermediate data by using the first type of feature data of the current user object held and the feature data of the initial first central object numbered i.
[0040] When specifically calculating the first intermediate data, the first server can, according to the arrangement order, subtract each feature in the feature data of the initial first central object from the corresponding feature in the first type of feature data of the current user object, and use the obtained difference as the first intermediate data corresponding to the feature.
[0041] For example, when calculating the first intermediate data of the feature numbered j (j is less than or equal to n), it can be calculated according to the following formula: <v j >1 = (<u j >1 – x j ). Among them, <v j >1 represents the first intermediate data numbered j, <u j >1 represents the feature numbered j in the feature data of the initial first central object numbered i, and x j represents the feature numbered j in the first type of feature data of the current user object. For features with numbers greater than n, since the first server does not have the corresponding feature data, 0 can be used instead. For example, when calculating the first intermediate data of the feature numbered q (q is greater than n and less than or equal to n + m), it can be calculated according to the following formula: <v q >1 = (<u q >1 – 0) = <u q >1. Among them, <v q >1 represents the first intermediate data numbered q, <u q >1 represents the feature numbered q in the feature data of the initial first central object numbered i. In the above manner, the first server can calculate n + m first intermediate data respectively: <v1> 1, <v2>1,...,<v t >1...,<v n+m >1. Among them, t represents an integer number greater than or equal to 1 and less than or equal to n + m.
[0042] Similarly, when specifically calculating the second intermediate data, the second server can, according to the arrangement order, subtract each feature in the feature data of the initial second central object from the corresponding feature in the second type of feature data of the current user object, and use the obtained difference as the second intermediate data corresponding to the feature.
[0043] For example, when calculating the second intermediate data of the feature numbered q (q is greater than n and less than or equal to n + m), it can be calculated according to the following formula: <v q >2 = (<u q >2 – y q ). Among them, <v q >2 represents the second intermediate data numbered q, <u q >2 represents the feature numbered q in the feature data of the initial second central object numbered i, and y q represents the feature numbered q in the second type of feature data of the current user object. For features numbered less than n, since the second server does not have corresponding feature data, 0 can be used instead. For example, when calculating the second intermediate data of the feature numbered j (j is less than or equal to n), it can be calculated according to the following formula: <v j >2 = (<u j >2 – 0) = <u j >2. Among them, <v j >2 represents the second intermediate data numbered j, <u j >2 represents the feature numbered j in the feature data of the initial second central object numbered i. In the above manner, the second server can calculate n + m second intermediate data respectively: <v1> 2, <v2>2,..., <v t >2..., <v n+m >2. Among them, t represents an integer number greater than or equal to 1 and less than or equal to n + m.
[0044] Then, the first server can, according to the preset protocol rules, use the first intermediate data (for example, <v t >1), and the square of the first intermediate data (for example, (<v t >1) 2 ) as inputs, and jointly perform a first function operation with the second data party (using the second intermediate data <v t >2, and the square of the second intermediate data (<v t >2) 2 as inputs), and specifically can calculate the feature distance between the current data object and the initial central object numbered i according to the following formula:
[0045]
[0046] Specifically in implementation, the first server and the second server can complete the above first function operation by adopting multi-party secure computation (MPC); and, enable the first server to obtain the first shard of the feature distance, and enable the second server to obtain the second shard of the feature distance.
[0047] In the above manner, the first server and the second server can, through cooperation, calculate the feature distances between each user object and each initial central object, and each obtain one shard of the feature distance.
[0048] Furthermore, the first server can, according to the preset protocol rules, use the first shard of the feature distance between the user object and the initial central object, and perform a second interaction with the second data party (using the second shard of the feature distance between the user object and the initial central object) to determine the initial central object closest to each user object as the matching central object of the user object.
[0049] The following takes determining the matching central object of the current user object as an example for specific description. The initial central objects include: initial central object 1, initial central object 2, and initial central object 3.
[0050] Among them, the first server holds the first shard d-1_1 of the feature distance between the current user object and the initial central object 1, the first shard d-2_1 of the feature distance between the current user object and the initial central object 2, and the first shard d-3_1 of the feature distance between the current user object and the initial central object 3. The second server holds the second shard d-1_2 of the feature distance between the current user object and the initial central object 1, the second shard d-2_2 of the feature distance between the current user object and the initial central object 2, and the second shard d-3_2 of the feature distance between the current user object and the initial central object 3.
[0051] First, the first server and the second server can first determine the existence of the following three distance relationship combinations. They are respectively denoted as: relationship combination 1 (including the two feature distances between the current user object and the initial central objects 1 and 2), relationship combination 2 (including the two feature distances between the current user object and the initial central objects 2 and 3), and relationship combination 3 (including the two feature distances between the current user object and the initial central objects 1 and 3).
[0052] Next, for relationship combination 1, according to the preset protocol rules, the first server can use the first shards of the two feature distances between the current user object and the initial central objects 1 and 2, and calculate the first shard of the intermediate data of relationship combination 1 according to the following formula: D-1_1 = d-1_1 - d-2_1.
[0053] Similarly, the second server can use the second shards of the two feature distances between the current user object and the initial central objects 1 and 2, and calculate the second shard of the intermediate data of relationship combination 1 according to the following formula: D-1_2 = d-1_2 - d-2_2.
[0054] Then, the first server can use the first shard of the intermediate data of relationship combination 1 as the input, and the second server (using the second shard of the intermediate data of relationship combination 1 as the input), and perform corresponding function (a comparison function) operations through multi-party secure computing to obtain the magnitude relationship of the two feature distances included in relationship combination 1. For example, the feature distance between the current user object and the initial central object 1 is greater than the feature distance between the current user object and the initial central object 2.
[0055] In a similar manner, the magnitude relationship of the two feature distances included in relationship combination 2 (for example, the feature distance between the current user object and the initial central object 2 is less than the feature distance between the current user object and the initial central object 3), and the magnitude relationship of the two feature distances included in relationship combination 3 (for example, the feature distance between the current user object and the initial central object 1 is greater than the feature distance between the current user object and the initial central object 3) can be obtained.
[0056] Furthermore, the first server can determine, based on the magnitude relationships of the feature distances included in the above three relationship combinations, that the feature distance between the current user object and the initial central object 2 is the smallest, and determine the initial central object 2 as the matching central object of the current user object.
[0057] In the above manner, the first server and the second server can cooperate to determine the matching central object of each user object.
[0058] The first server and / or the second server can divide the user objects into clusters corresponding to the matching central objects.
[0059] Specifically, for example, the matching central object of user object 1 is the initial central object 1. Therefore, user object 1 can be divided into the cluster corresponding to the initial central object 1 (which can be denoted as the first cluster). Another example, the matching central object of user object 2 is the initial central object 3. Therefore, user object 2 can be divided into the cluster corresponding to the initial central object 3 (which can be denoted as the third cluster).
[0060] Through the above manner, the first server and the second server can cooperate to complete the initial clustering. The first server can cluster to obtain 3 clusters, namely: the first cluster, the second cluster, and the third cluster. Among them, each cluster contains one or more user objects with similar purchase potentials.
[0061] Next, the first server and the second server can, according to the preset protocol rules, respectively use the first type of feature data and the second type of feature data of the user objects in each cluster they respectively hold to perform a third interaction to determine the cluster centers in each cluster as the first central objects.
[0062] Specifically, taking the determination of the first central object of the first cluster as an example, a detailed description is given.
[0063] The above first cluster may contain a total of two user objects, namely user object 1 and user object 3. Among them, the first type of feature data of user object 1 held by the first server is denoted as: x1 = (x11, x12, x13,..., x1 n ), and the first type of feature data of user object 3 is denoted as: x3 = (x31, x32, x33,..., x3 n ). The second type of feature data of user object 1 held by the second server is denoted as: y1 = (y11, y12, y13,..., y1 m ), and the second type of feature data of user object 3 is denoted as: x3 = (y31, y32, y33,..., y3 m ).
[0064] The first server can take x1 and x3 as inputs and cooperate with the second server (taking y1 and y3 as inputs) to perform corresponding function operations using multi-party secure computing. By calculating the average value of each feature, the average value of the feature data of the user objects in the first group can be determined, which can be denoted as: Z’ = ((x11 + x31) / 2, (x12 + x32) / 2,..., (x1 n + x3 n ) / 2), (y11 + y31) / 2, (y12 + y32) / 2,..., (y1 m + y3 m ) / 2). Furthermore, Z’ can be used as the feature data of the first central object of the first group.
[0065] In specific implementation, according to the preset protocol rules, through the above function operations, the first server can only obtain the first type of feature data of the first central object of the first group: ((x11 + x31) / 2, (x12 + x32) / 2,..., (x1 n + x3 n ) / 2)); the second server can only obtain the second type of feature data of the first central object of the first group: ((y11 + y31) / 2, (y12 + y32) / 2,..., (y1 m + y3 m ) / 2).
[0066] Of course, in some cases, the first server and / or the second server can also obtain the complete feature data Z’ of the first central object of the first group.
[0067] In the above manner, the first server and the second server can cooperate to determine the first central object of the first group, the first central object of the second group, and the first central object of the third group.
[0068] Furthermore, the first server and the second server can cooperate to detect whether the currently clustered first group, second group, and third group meet the requirements.
[0069] Specifically, the first server can, according to the preset protocol rules, use the feature data of the initial first central object and the first type of feature data of the first central object, and cooperate with the second server (using the feature data of the initial second central object and the second type of feature data of the first central object) to calculate the feature distance between the initial central object and the first central object through multi-party secure computing.
[0070] For example, the first server and the second server can cooperate in the above manner to calculate the feature distance between the initial central object corresponding to the first cluster (i.e., the combination of the initial first central object and the initial second central object) and the first central object, and determine whether the feature distance is less than a preset distance threshold. In the case where it is determined that the above feature distance is less than the preset distance threshold, it can be determined whether the first cluster meets the requirements. On the contrary, in the case where it is determined that the above feature distance is greater than or equal to the preset distance threshold, it can be determined that the first cluster does not meet the requirements.
[0071] In the above manner, if it is determined that all of the above three clusters meet the requirements, it can be determined that the clustering has been completed, and the above three clusters are determined as the final clustering result.
[0072] On the contrary, if it is determined that at least one of the above three clusters does not meet the requirements, it can be determined that further clustering is still needed, and then it can be triggered to use the newly determined first central object to replace the initial central object for the next clustering.
[0073] Specifically, the first server can, according to the preset protocol rules, use the first type of feature data of the user object and the first type of feature data of the first central object to cooperate with the second server (which uses the second type of feature data of the same user object and the second type of feature data of the same first central object) to perform the first interaction and calculate the first shard of the feature distance between the user object and the first central object. At the same time, the second server can calculate the second shard of the feature distance between the user object and the second central object. The above data processing process can specifically refer to the data processing process of the shards when the first server and the second server cooperate to calculate the feature distance between the user object and the initial central object. Here, it will not be elaborated.
[0074] Then, the first server can, according to the preset protocol rules, use the first shard of the feature distance between the user object and the first central object to cooperate with the second server (which uses the second shard of the feature distance between the same user object and the same first central object) to perform the second interaction to determine the matching central object of each user object from the first central object. The above data processing process can specifically refer to the data processing process when the first server and the second server cooperate to determine the matching central object of the user object from multiple initial central objects. Here, it will not be elaborated.
[0075] Furthermore, the first server re-clusters the user objects, and divides the user objects into the clusters corresponding to the above matching central objects respectively, obtaining three new clusters.
[0076] Further, the first server and the second server can, according to the preset protocol rules, respectively calculate the current new cluster center as the second center object by using the first type of feature data and the second type of feature data of the data objects in the currently new cluster they hold. Then, according to the preset protocol rules, by calculating and detecting whether the feature distance between the feature data of the first center object and the second center object is less than the preset distance threshold, it is determined whether the currently new cluster meets the requirements. The above data processing process can specifically refer to the data processing process of the first server and the second server cooperating to detect whether the three clusters of the initial clustering meet the requirements. Regarding this, no further elaboration will be made.
[0077] In the case where it is determined that the currently new cluster meets the requirements, it can be determined that the clustering for the user object is completed, and the currently new cluster is used as the final clustering result.
[0078] On the contrary, in the case where it is determined that the currently new cluster does not meet the requirements, clustering needs to continue. Specifically, the second center object newly determined last time can be used as the first center object for the current time, and the above data processing process is repeated until a cluster that meets the requirements is obtained.
[0079] In the above manner, the first server and the second server can complete the clustering process for the user object through cooperation. And, according to the relevant protocol, the first server can obtain the final clustering result.
[0080] The first server can determine the user objects in the first cluster, the user objects in the second cluster, and the user objects in the third cluster according to the final clustering result.
[0081] Further, the first server can establish a product push strategy for the user objects in the first cluster, denoted as the first push strategy, by obtaining and learning the historical shopping records of the user objects in the first cluster. At the same time, by obtaining and learning the historical shopping records of the user objects in the second cluster, establish a product push strategy for the user objects in the second cluster, denoted as the second push strategy; obtain and learn the historical shopping records of the user objects in the third cluster, and establish a product push strategy for the user objects in the third cluster, denoted as the third push strategy.
[0082] Furthermore, the first server can perform product push for user objects in different clusters by using the corresponding push strategies to obtain a better push effect.
[0083] Specifically, for example, a shopping website plans to push products to target user objects to increase the conversion rate of the target user objects on the shopping website. The first server can determine that the target user object belongs to the second group by retrieving the identity identifiers of user objects in the first group, the second group, and the third group according to the identity identifier of the target user object. Furthermore, according to the second push strategy corresponding to the second group, the target product suitable for the target user object and the target push method can be determined. Further, a target push message related to the target product can be generated and pushed to the target user object in the target push method. Thus, the conversion rate of the target user object when shopping on the shopping website can be effectively increased.
[0084] Combined Figure 2 with Figure 3 As shown, an embodiment of this specification provides a clustering method for data objects. When this method is specifically implemented, it may include the following content.
[0085] S201: According to the preset protocol rules, use the first type of feature data of the data object held and the first type of feature data of the first central object to perform a first interaction with the second data party to obtain a first shard of the feature distance between the data object and the first central object; wherein, the first central object is the class group center determined in the previous clustering; the second data party holds the second type of feature data of the data object and the second type of feature data of the first central object; the second data party obtains a second shard of the feature distance between the data object and the first central object by performing a first interaction with the first data party.
[0086] S202: According to the preset protocol rules, use the first shard of the feature distance between the data object and the first central object to perform a second interaction with the second data party to determine the matching central object of the data object; wherein, the matching central object is the first central object with the smallest feature distance from the data object.
[0087] S203: Divide the data object into the class group corresponding to the matching central object.
[0088] In some embodiments, the above preset protocol rules may specifically be a protocol rule based on secret sharing and oblivious transfer.
[0089] Among them, oblivious transfer (OT) may specifically refer to a communication protocol that can protect the data privacy of both parties. Based on this protocol, neither communication party can know the specific data input by the other party, enabling the communication parties to transmit data information in a way of selective obfuscation.
[0090] Secret Sharing (SS), also known as secret sharing, can specifically refer to a multi-party security protocol. Based on secret sharing, a secret can be split in an appropriate way. Each split share (denoted as a share) can be held and managed by different participants, and a single participant cannot recover the secret information based on a single share. Only several participants collaborating together can recover the secret message.
[0091] By interacting through the above preset protocol rules, the first server and the second server can securely complete various related function operations without knowing the specific data input by each other.
[0092] In some embodiments, the method can specifically be applied to a first data party. Among them, the first data party holds and manages the first type of feature data of a data object. The second data party holds and manages the second type of feature data of the same data object.
[0093] In some embodiments, the above data object can specifically be a user object, an enterprise object, a commodity object, etc. Of course, the above-listed data objects are only illustrative. In specific implementation, according to the specific application scenario and processing requirements, the above data object can also include other types of data objects. This specification does not make any limitations in this regard.
[0094] In some embodiments, the above first type of feature data and the second type of feature data can specifically be data for describing different types of attribute features of the same data object.
[0095] Specifically, taking the user object as an example. The above first type of feature data can specifically be the shopping-related features of the user object. For example, the shopping frequency of the user object, the monthly shopping consumption amount of the user object, the commodity types of the commodities purchased by the user object, the commodity brands most frequently purchased by the user object, etc.
[0096] The above second type of feature data can specifically be the asset-related features of the user object. For example, the monthly income of the user object, the mortgage record of the user object, the financial management income of the user object, etc.
[0097] Of course, it should be noted that the above-listed first type of feature data and the second type of feature data are only illustrative. In specific implementation, according to the specific application scenario and processing requirements, the above first type of feature data and the second type of feature data can specifically also include other types and contents of feature data. This specification does not make any limitations in this regard.
[0098] In some embodiments, the first type of feature data of the data object held by the first data party and the second type of feature data of the data object held by the second data party can specifically be pre-aligned.
[0099] Therefore, among the first type of feature data of the data objects held by the first data party and the second type of feature data of the data objects held by the second data party, the feature data sorted at the same position or with the same corresponding object identifier belong to the first type of feature data and the second type of feature data of the same data object.
[0100] In some embodiments, the above-mentioned first type of feature data may specifically include multiple first features. Among them, the multiple first features may be specifically arranged in a preset first order. Each first type of feature data corresponds to the object identifier of a data object.
[0101] Similarly, the above-mentioned second type of feature data includes multiple second features. Among them, the multiple second features may be specifically arranged in a preset second order. Each second type of feature data corresponds to the object identifier of a data object.
[0102] The first type of feature data of the data objects held by the first data party and the second type of feature data of the data objects held by the second data party can be combined to obtain the complete feature data of the data object.
[0103] In some embodiments, the above-mentioned first central object can be specifically understood as the cluster center of the cluster determined in the previous clustering. Specifically, the first data party may hold the first type of feature data of the first central object, and the second data party may hold the second type of feature data of the first central object.
[0104] Of course, in some cases, the first data party may also hold the complete feature data of the first central object. The second data party may also hold the complete feature data of the first central object.
[0105] In some embodiments, according to the preset protocol rules, using the first type of feature data of the held data object and the first type of feature data of the first central object, a first interaction is performed with the second data party to obtain a first shard of the feature distance between the data object and the first central object. When specifically implemented, it may include the following content: using the first type of feature data of the held data object and the first type of feature data of the first central object, calculating the first intermediate data and the square of the first intermediate data; using the first intermediate data and the square of the first intermediate data as inputs, performing a first function operation with the second data party (using the second intermediate data and the square of the second intermediate data as inputs) to obtain a first shard of the feature distance between the data object and the first central object. At the same time, the second data party can obtain a second shard of the feature distance.
[0106] In some embodiments, the first intermediate data may specifically be the difference between the corresponding features of the first type of feature data of the data object and the first type of feature data of the first central object. The second intermediate data may specifically be the difference between the corresponding features of the second type of feature data of the data object and the second type of feature data of the second central object.
[0107] In some embodiments, the first function operation may specifically be understood as a function operation designed based on the Secure Multi-Party Computation (MPC) protocol.
[0108] When specifically performing the first function operation, the feature distance between the data object and the first central object may be calculated based on the input first intermediate data, the square of the first intermediate data, the second intermediate data, and the square of the second intermediate data, and the first shard of the feature distance may be output to the first server, and the second shard of the feature distance may be output to the second server. Among them, the first shard of the feature distance and the second shard of the feature distance can be combined to recover the complete feature distance.
[0109] In some embodiments, the feature distance may specifically include a Euclidean distance calculated based on a feature vector.
[0110] In some embodiments, according to the preset protocol rules, the first shard of the feature distance between the data object and the first central object is used to perform a second interaction with the second data party to determine the matching central object of the data object. Specifically, when implemented, it may include the following content: Determine the matching central object of the current data object in the following manner: Determine a plurality of distance relationship combinations; where a distance relationship combination includes the feature distances between the current data object and two different first central objects; Use the first shard of the feature distance between the current data object and the first central object to calculate the first shard of the intermediate data of the distance relationship combination; According to the preset protocol rules, use the first shard of the intermediate data of the distance relationship combination to interact with the second data party to determine the magnitude relationship of the feature distances in the distance relationship group; According to the magnitude relationship of the feature distances, determine the first central object with the smallest feature distance from the current data object as the matching central object of the current data object.
[0111] Through the above embodiments, the first data party and the second data party can cooperate according to the preset protocol rules to determine the magnitude relationship between the feature distances between the current data object and the first central object pairwise (corresponding to the magnitude relationship of the feature distances in a distance relationship combination); furthermore, according to the magnitude relationship between the feature distances pairwise, the first central object closest to the current data object can be determined as the matching central object of the current data object.
[0112] Referring to the above method for determining the matching central object of the current data object, the matching central objects of each data object can be determined respectively.
[0113] In some embodiments, during specific implementation, refer to Figure 3 , each data object can be classified into the group corresponding to its matching central object according to the matching central objects of each data object, thus completing one clustering and obtaining multiple groups.
[0114] Among them, each of the multiple groups corresponds to a first central object. Each group can specifically include one or more data objects with similar features.
[0115] In some embodiments, before the first interaction with the second data party according to the preset protocol rules, using the first type of feature data of the held data object and the first type of feature data of the first central object, when the method is specifically implemented, the following content can also be included: responding to the clustering request for the data object, determining the number of groups for clustering; randomly generating the feature data of multiple initial first central objects according to the number of groups; where the second data party randomly generates the feature data of multiple initial second central objects; according to the preset protocol rules, using the first type of feature data of the held data object and the feature data of the initial first central object, interacting with the second data party (using the second type of feature data of the held data object and the feature data of the initial second central object) for the fifth time to obtain the first shard of the feature distance between the data object and the initial central object. Among them, the initial central object is obtained by combining the initial first central object and the initial second central object; the second data party obtains the second shard of the feature distance between the data object and the initial central object by interacting with the first data party for the fifth time.
[0116] Further, the first data party can, according to the preset protocol rules, use the first shard of the feature distance between the data object and the initial central object to interact with the second data party (using the second shard of the feature distance between the data object and the initial central object) for the second time to determine the initial matching central object of the data object. Among them, the initial matching central object can specifically be the initial central object with the smallest feature distance from the data object.
[0117] Then, the first data party can classify the data object into the group corresponding to the initial matching central object.
[0118] Among them, the specific value of the number of groups for the above clustering can be flexibly set according to specific circumstances. This specification does not make any limitation on this. The above clustering request for the data object can specifically be initiated by the first data party or by the second data party.
[0119] In some embodiments, the above-mentioned initial central object can specifically be understood as a central object obtained by combining the initial first central object and the initial second central object.
[0120] In some embodiments, referring to Figure 3 , after dividing the data objects into the clusters corresponding to the matching central objects, when the method is specifically implemented, the following content may further be included: determining the data objects in the current cluster; according to a preset protocol rule, using the first type of feature data of the data objects in the current cluster to perform a third interaction with a second data party (using the second type of feature data of the data objects in the current cluster), determining the cluster center of the current cluster as the second central object, and obtaining the first type of feature data of the second central object; wherein, the second data party obtains the second type of feature data of the second central object by performing a third interaction with the first data party.
[0121] Among them, the above-mentioned second central object can specifically be understood as the cluster center of the cluster determined in the current clustering. Specifically, the first data party may hold the first type of feature data of the second central object, and the second data party may hold the second type of feature data of the second central object.
[0122] Of course, in some cases, the first data party may also hold the complete feature data of the second central object. The second data party may also hold the complete feature data of the second central object.
[0123] In some embodiments, the above-mentioned using the first type of feature data of the data objects in the current cluster to perform a third interaction with a second data party (using the second type of feature data of the data objects in the current cluster), determining the cluster center of the current cluster as the second central object, when specifically implemented, may include the following content: using the second type of feature data of the data objects in the current cluster as input to perform a second function operation with the second data party to calculate the average value of the feature data of the data objects in the current cluster, and determining the cluster center of the current cluster as the second central object based on the average value of the feature data. Among them, the above-mentioned second function operation may specifically be a function operation based on a multi-party secure computing protocol.
[0124] When specifically performing the above-mentioned second function operation, the average value of the feature data of the data objects in the current cluster may be obtained by respectively calculating the average value of each feature according to the input first type of feature data of the data objects in the current cluster and the second type of feature data of the data objects in the current cluster.
[0125] The data object indicated by the average value of the feature data of the data objects in the current cluster may be used as the cluster center of the current cluster, that is, the second central object corresponding to the current cluster.
[0126] In some embodiments, after determining the cluster center of the current cluster as the second central object and obtaining the first type of feature data of the second central object, refer to Figure 3 , when the method is specifically implemented, the following content may further be included: According to the preset protocol rules, use the first type of feature data of the second central object, the first type of feature data of the first central object, and perform a fourth interaction with the second data party to detect whether the feature distance between the second central object and the first central object is less than a preset distance threshold; in the case where it is determined that the feature distance between the second central object and the first central object is less than the preset distance threshold, determine that the current cluster meets the requirements.
[0127] Through the above embodiments, it is possible to judge whether the current cluster meets the requirements by calculating and based on the feature distance between the second central object of the current cluster obtained by the current clustering and the first central object of the cluster obtained by the previous clustering.
[0128] In the case where it is determined that all current clusters meet the requirements, clustering can be stopped. At this time, the multiple clusters obtained by the current clustering can be determined as the final clustering result. The clustering of the data objects is completed.
[0129] On the contrary, in the case where it is determined that at least one of the current clusters does not meet the requirements, clustering needs to be continued. At this time, the second central object determined by the current clustering can be used as the first central object, and the above clustering process is repeated until the feature distance between the second central object of the newly clustered cluster and the first central object of the previous time is less than the preset distance threshold, and the final clustering result is obtained. The clustering of the data objects is completed.
[0130] In some embodiments, when specifically implemented, data objects in each cluster can be determined according to the final clustering result. Then, according to the characteristics of the data objects in each cluster, a preset push strategy matching the cluster is constructed. And a corresponding relationship between each cluster and the preset push strategy that matches it is established.
[0131] In some embodiments, when data needs to be pushed to a target data object, the target cluster to which the target data object belongs can be determined first. Then, the preset push strategy corresponding to the target cluster is found as the target push strategy. Further, according to the target push strategy, the target data suitable for the target data object and the target push method can be determined. Furthermore, the target data can be pushed to the target data object in the target push method. Thus, a better push effect can be obtained.
[0132] As can be seen from the above, based on the clustering method of data objects provided in the embodiments of this specification, when the first data party holding the first type of feature data of the data object and the second data party holding the second type of feature data of the data object cooperate to cluster the data object, they can first perform a first interaction using the data they each hold according to the preset protocol rules, respectively, to obtain a shard of the feature distance between the data object and the first central object; then, according to the preset protocol rules, they can perform a second interaction using the shards of the feature distance they each hold, respectively, to determine the matching central object for each data object from the first data objects; and then divide the data objects into the clusters corresponding to the matching central objects to complete one clustering of the data objects. Thus, it can enable the first data party and the second data party to successfully complete the clustering of the data objects through cooperation without disclosing the data they each hold to the other party, effectively protecting the data privacy of both parties and reducing the risk of the data they each hold being leaked during the clustering process.
[0133] The embodiments of this specification also provide a server, including a processor and a memory for storing processor-executable instructions. When specifically implemented, the processor can execute the following steps according to the instructions: According to the preset protocol rules, use the first type of feature data of the data object held and the first type of feature data of the first central object to perform a first interaction with the second data party to obtain a first shard of the feature distance between the data object and the first central object; where the first central object is the cluster center determined in the previous clustering; the second data party holds the second type of feature data of the data object and the second type of feature data of the first central object; the second data party obtains a second shard of the feature distance between the data object and the first central object by performing a first interaction with the first data party; according to the preset protocol rules, use the first shard of the feature distance between the data object and the first central object to perform a second interaction with the second data party to determine the matching central object of the data object; where the matching central object is the first central object with the smallest feature distance from the data object; divide the data object into the cluster corresponding to the matching central object.
[0134] In order to be able to complete the above instructions more accurately, refer to Figure 4 As shown, the embodiments of this specification also provide another specific server. Among them, the server includes a network communication port 401, a processor 402, and a memory 403. The above structures are connected by internal cables so that each structure can perform specific data interactions.
[0135] Among them, the network communication port 401 can specifically be used to initiate a clustering request for the data object.
[0136] The processor 402 can be specifically configured to perform a first interaction with a second data party according to a preset protocol rule, using the first type of feature data of the data object it holds and the first type of feature data of the first central object, so as to obtain a first shard of the feature distance between the data object and the first central object. Wherein, the first central object is the cluster center determined in the previous clustering; the second data party holds the second type of feature data of the data object and the second type of feature data of the first central object; the second data party obtains a second shard of the feature distance between the data object and the first central object by performing a first interaction with a first data party; according to the preset protocol rule, perform a second interaction with the second data party using the first shard of the feature distance between the data object and the first central object, so as to determine a matching central object of the data object. Wherein, the matching central object is the first central object with the smallest feature distance from the data object; divide the data object into the cluster corresponding to the matching central object.
[0137] The memory 403 can be specifically configured to store corresponding instruction programs.
[0138] In this embodiment, the network communication port 401 can be bound to different communication protocols, so as to send or receive different data virtual ports. For example, the network communication port can be a port responsible for web data communication, or a port responsible for FTP data communication, or a port responsible for email data communication. In addition, the network communication port can also be a physical communication interface or communication chip. For example, it can be a wireless mobile network communication chip, such as GSM, CDMA, etc.; it can also be a Wifi chip; it can also be a Bluetooth chip.
[0139] In this embodiment, the processor 402 can be implemented in any suitable manner. For example, the processor can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, application specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers, etc. This specification does not make any limitations.
[0140] In this embodiment, the memory 403 can include multiple levels. In a digital system, anything that can store binary data can be a memory; in an integrated circuit, a circuit with a storage function without a physical form is also called a memory, such as RAM, FIFO, etc.; in a system, a storage device with a physical form is also called a memory, such as a memory stick, TF card, etc.
[0141] The embodiment of this specification also provides a computer-readable storage medium based on the above clustering method for data objects. The computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed, it realizes: according to the preset protocol rules, using the first type of feature data of the held data object and the first type of feature data of the first central object, to perform a first interaction with the second data party to obtain a first shard of the feature distance between the data object and the first central object; wherein, the first central object is the class group center determined in the previous clustering; the second data party holds the second type of feature data of the data object and the second type of feature data of the first central object; the second data party obtains a second shard of the feature distance between the data object and the first central object by performing a first interaction with the first data party; according to the preset protocol rules, using the first shard of the feature distance between the data object and the first central object, to perform a second interaction with the second data party to determine the matching central object of the data object; wherein, the matching central object is the first central object with the smallest feature distance from the data object; and dividing the data object into the class group corresponding to the matching central object.
[0142] In this embodiment, the above storage medium includes but is not limited to Random Access Memory (RAM), Read-Only Memory (ROM), Cache, Hard Disk Drive (HDD), or Memory Card. The memory can be used to store computer program instructions. The network communication unit can be set according to the standards stipulated by the communication protocol and is used for the interface of network connection communication.
[0143] In this embodiment, the functions and effects specifically realized by the program instructions stored in this computer-readable storage medium can be explained by comparison with other embodiments and will not be elaborated here.
[0144] Refer to Figure 5 As shown, at the software level, the embodiment of this specification also provides a clustering device for data objects, and this device can specifically include the following structural modules:
[0145] The first interaction module 501 can be specifically used to perform a first interaction with a second data party according to a preset protocol rule, using the first type of feature data of the data object it holds and the first type of feature data of the first central object, so as to obtain a first fragment of the feature distance between the data object and the first central object; wherein, the first central object is the cluster center determined in the previous clustering; the second data party holds the second type of feature data of the data object and the second type of feature data of the first central object; the second data party obtains a second fragment of the feature distance between the data object and the first central object by performing a first interaction with the first data party.
[0146] The second interaction module 502 can be specifically used to perform a second interaction with the second data party according to a preset protocol rule, using the first fragment of the feature distance between the data object and the first central object, so as to determine the matching central object of the data object; wherein, the matching central object is the first central object with the smallest feature distance from the data object.
[0147] The classification module 503 can be specifically used to classify the data object into the cluster corresponding to the matching central object.
[0148] It should be noted that the units, devices or modules etc. described in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. For the convenience of description, when describing the above devices, they are divided into various modules according to functions for separate description. Of course, when implementing this specification, the functions of each module can be implemented in the same or multiple software and / or hardware, or the modules implementing the same function can be realized by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in electrical, mechanical or other forms.
[0149] As can be seen from the above, based on the clustering device for data objects provided in the embodiments of this specification, the first data party and the second data party can cooperate to complete the clustering of data objects without disclosing the data they hold to each other, protecting the data privacy of both parties and reducing the risk of the data they hold being leaked during the clustering process.
[0150] Although this specification provides method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way among the execution orders of numerous steps and does not represent the only execution order. When the actual device or client product is executed, it may be executed in the order of the method shown in the embodiments or the drawings or executed in parallel (for example, in an environment of parallel processors or multi-threaded processing, or even in a distributed data processing environment). The terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, product or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, product or device. Without further limitation, there is no exclusion of additional identical or equivalent elements in the process, method, product or device comprising the said elements. The terms such as first, second, etc. are used to denote names and do not denote any particular order.
[0151] Those skilled in the art also know that in addition to implementing the controller in the form of pure computer-readable program code, the method steps can be logically programmed to enable the controller to implement the same functions in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. Therefore, such a controller can be regarded as a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or structures within the hardware component.
[0152] This specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc. that perform specific tasks or implement specific abstract data types. This specification can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer-readable storage media including storage devices.
[0153] As can be seen from the description of the above embodiments, those skilled in the art can clearly understand that this specification can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of this specification can essentially be embodied in the form of a software product, and this computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to enable a computer device (which can be a personal computer, mobile terminal, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments of this specification.
[0154] The embodiments in this specification are described in a progressive manner. For the same or similar parts between the embodiments, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. This specification can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on.
[0155] Although this specification is depicted through embodiments, those of ordinary skill in the art know that this specification has many variations and changes without departing from the spirit of this specification. It is hoped that the appended claims will cover these variations and changes without departing from the spirit of this specification. < / v1> < / v1> < / u2> < / u2> < / u1> < / u1> < / u1> < / u1>
Claims
1. A clustering method for data objects, comprising: According to the preset protocol rules, use the first type of feature data of the data object held and the first type of feature data of the first central object to perform a first interaction with the second data party to obtain a first shard of the feature distance between the data object and the first central object; wherein, the first central object is the cluster center determined by the previous clustering; the second data party holds the second type of feature data of the data object and the second type of feature data of the first central object; the second data party obtains a second shard of the feature distance between the data object and the first central object by performing a first interaction with the first data party; the data object includes a user object; the first type of feature data is the shopping-related feature of the user object, including at least one of the following: the shopping frequency of the user object, the monthly shopping consumption amount of the user object, the type of goods purchased by the user object; the second type of feature is the asset-related feature of the user object, including at least one of the following: the monthly income of the user object, the mortgage record of the user object, the financial management income of the user object. According to the preset protocol rules, use the first shard of the feature distance between the data object and the first central object to perform a second interaction with the second data party to determine the matching central object of the data object; including: determining the matching central object of the current data object in the following manner: determining a plurality of distance relationship combinations; wherein, one distance relationship combination includes the feature distances between the current data object and two different first central objects; use the first shard of the feature distance between the current data object and the first central object to calculate the first shard of the intermediate data of the distance relationship combination; according to the preset protocol rules, use the first shard of the intermediate data of the distance relationship combination to perform an interaction with the second data party to determine the magnitude relationship of the feature distances in the distance relationship group; according to the magnitude relationship of the feature distances, determine the first central object with the smallest feature distance from the current data object as the matching central object of the current data object; wherein, the matching central object is the first central object with the smallest feature distance from the data object. Divide the data object into the cluster corresponding to the matching central object.
2. The method according to claim 1, after dividing the data objects into the clusters corresponding to the matching central objects, the method further comprises: Determine the data objects in the current cluster. According to the preset protocol rules, use the first type of feature data of the data objects in the current cluster to perform a third interaction with the second data party to determine the cluster center of the current cluster as the second central object and obtain the first type of feature data of the second central object; wherein, the second data party obtains the second type of feature data of the second central object by performing a third interaction with the first data party.
3. The method according to claim 2, after determining the cluster center of the current cluster as the second central object and obtaining the first type of feature data of the second central object, the method further comprises: According to the preset protocol rules, use the first type of feature data of the second central object and the first type of feature data of the first central object to perform a fourth interaction with the second data party to detect whether the feature distance between the second central object and the first central object is less than a preset distance threshold. In the case where it is determined that the feature distance between the second central object and the first central object is less than the preset distance threshold, determine that the current cluster meets the requirements.
4. The method according to claim 1, before performing a first interaction with a second data party according to a preset protocol rule, using the first type of feature data of the held data objects and the first type of feature data of the first central object, the method further comprises: Respond to the clustering request for the data object and determine the number of clusters for clustering. Generate the characteristic data of multiple initial first central objects randomly according to the number of the clusters; wherein, the second data party generates the characteristic data of multiple initial second central objects randomly. According to the preset protocol rules, use the first type of characteristic data of the data objects held, and the characteristic data of the initial first central objects, to perform a fifth interaction with the second data party, so as to obtain the first shard of the characteristic distance between the data objects and the initial central objects; wherein, the initial central objects are obtained by combining the initial first central objects and the initial second central objects; the second data party obtains the second shard of the characteristic distance between the data objects and the initial central objects by performing a fifth interaction with the first data party. According to the preset protocol rules, use the first shard of the characteristic distance between the data objects and the initial central objects, to perform a second interaction with the second data party, and determine the initial matching central objects of the data objects; wherein, the initial matching central objects are the initial central objects with the smallest characteristic distance from the data objects. Divide the data objects into the clusters corresponding to the initial matching central objects.
5. The method according to claim 1, performing a first interaction with a second data party according to a preset protocol rule, using the first type of feature data of the held data objects and the first type of feature data of the first central object, to obtain a first shard of the feature distance between the data objects and the first central object, comprising: Use the first type of characteristic data of the data objects held, and the first type of characteristic data of the first central objects, to calculate the first intermediate data and the square of the first intermediate data. Use the first intermediate data and the square of the first intermediate data as inputs, and perform a first function operation with the second data party, so as to obtain the first shard of the characteristic distance between the data objects and the first central objects.
6. The method according to claim 2, performing a third interaction with a second data party according to a preset protocol rule, using the first type of feature data of the data objects in the current cluster, to determine the cluster center of the current cluster as the second central object, comprising: Use the second type of characteristic data of the data objects in the current cluster as an input, and perform a second function operation with the second data party, so as to calculate the average value of the characteristic data of the data objects in the current cluster, and determine the cluster center of the current cluster as the second central object based on the average value of the characteristic data.
7. The method according to claim 1, the preset protocol rule includes a protocol rule based on secret sharing and oblivious transfer.
8. The method according to claim 1, the data objects include user objects; Correspondingly, the first type of feature data includes shopping features of the user object; the second type of feature data includes asset features of the user object.
9. A clustering device for data objects, comprising: A first interaction module, configured to, according to the preset protocol rules, use the first type of characteristic data of the data objects held, and the first type of characteristic data of the first central objects, to perform a first interaction with the second data party, so as to obtain the first shard of the characteristic distance between the data objects and the first central objects; wherein, the first central objects are the cluster centers determined in the previous clustering; the second data party holds the second type of characteristic data of the data objects, and the second type of characteristic data of the first central objects; the second data party obtains the second shard of the characteristic distance between the data objects and the first central objects by performing a first interaction with the first data party; the data objects include user objects; the first type of characteristic data is the shopping type of characteristic of the user objects, including at least one of the following: the shopping frequency of the user objects, the monthly shopping consumption amount of the user objects, the types of commodities purchased by the user objects; the second type of characteristic is the asset type of characteristic of the user objects, including at least one of the following: the monthly income of the user objects, the mortgage record of the user objects, the financial management income of the user objects. The second interaction module is used to perform a second interaction with a second data party by using a first shard of the feature distance between a data object and a first central object according to a preset protocol rule, so as to determine a matching central object of the data object; wherein, the matching central object is the first central object with the smallest feature distance from the data object; wherein, the second interaction module determines the matching central object of the current data object in the following manner: determining a plurality of distance relationship combinations; wherein, a distance relationship combination includes the feature distances between the current data object and two different first central objects; calculating a first shard of the intermediate data of the distance relationship combination by using the first shard of the feature distance between the current data object and the first central object; according to the preset protocol rule, performing an interaction with the second data party by using the first shard of the intermediate data of the distance relationship combination to determine the magnitude relationship of the feature distances in the distance relationship group; according to the magnitude relationship of the feature distances, determining the first central object with the smallest feature distance from the current data object as the matching central object of the current data object; The classification module is used to classify the data object into the class group corresponding to the matching central object.
10. A server, comprising a processor and a memory for storing processor-executable instructions, wherein when the processor executes the instructions, the steps of the method according to any one of claims 1 to 7 are implemented.
11. A computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a computer device, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Method and device for clustering private data of multiple parties
CN111523143A