Clustering method for samples, server, client, device and readable storage medium

By receiving and processing sample sequence number and distance information on the server, determining and sending clustering results, the client can complete sample clustering, solving the data leakage problem that exists during clustering on the server and realizing data privacy protection.

CN112215289BActive Publication Date: 2025-06-03WEBANK (CHINA)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011095822.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-14
Publication Date
2025-06-03
Estimated Expiration
2040-10-14

AI Technical Summary

Technical Problem

There is a problem of data leakage when clustering samples, especially when processing samples containing user or enterprise private information.

Method used

The server receives the sample sequence number and distance information sent by the client, determines the first sample sequence number corresponding to each second sample sequence number, and sends it back to the client, so that the client determines the normal sample and the central sample as the same cluster samples, thereby avoiding the server from obtaining real sample data.

Benefits of technology

It realizes clustering of samples without leaking client data, enhances data privacy protection and avoids the risk of data leakage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112215289B_ABST
    Figure CN112215289B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of fintech, and discloses a clustering method for samples, a server, a client, a device, and a readable storage medium. The clustering method for samples includes: the server receives the sample numbers of the samples in the data set sent by each client and the distances corresponding to the sample numbers, where the distance is the distance between the normal sample corresponding to the sample number and the central sample in the data set; determines each first sample number corresponding to each second sample number according to each distance, the sample numbers include the first sample number of the normal sample and the second sample number of the central sample, and the data set includes multiple central samples; sends each first sample number corresponding to each second sample number to each client respectively, wherein the client determines the normal sample corresponding to the second sample number and the central sample as samples in the same cluster. The present invention avoids the leakage of data during the sample clustering of the client.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of financial technology (Fintech), and particularly to a clustering method for samples, a server, a client, a device, and a readable storage medium. Background Art

[0002] With the development of computer technology, more and more technologies are applied in the financial field. The traditional financial industry is gradually transforming into financial technology (Fintech). However, due to the security and real-time requirements of the financial industry, higher requirements are also put forward for technologies.

[0003] The clustering of samples is widely used in unsupervised machine learning. Currently, the server can cluster the samples of multiple clients. However, the samples carry private information of users or enterprises, resulting in the problem of data leakage when the server clusters the samples. Summary of the Invention

[0004] The main object of the present invention is to provide a clustering method for samples, a server, a client, a device, and a readable storage medium, aiming to solve the problem of data leakage when the server clusters the samples.

[0005] To achieve the above object, the present invention provides a clustering method for samples, and the clustering method for samples includes:

[0006] The server receives the sample numbers of the samples in the data set sent by each client and the distances corresponding to the sample numbers, where the distance is the distance between the normal sample corresponding to the sample number and the central sample in the data set;

[0007] Determine each first sample number corresponding to each second sample number according to each of the distances, where the sample numbers include the first sample number of the normal sample and the second sample number of the central sample, and the data set includes multiple central samples;

[0008] Send each first sample number corresponding to each second sample number to each of the clients, where the client determines the normal sample and the central sample corresponding to the second sample number as the samples in the same cluster.

[0009] In an embodiment, the step of determining each first sample number corresponding to each second sample number according to each of the distances includes:

[0010] Determine a target distance among each of the distances, and determine the sum of the distances of each of the target distances, where the first sample numbers and the second sample numbers corresponding to each of the target distances are the same;

[0011] Among the respective distance sums corresponding to the same first sample serial number, determine the minimum distance sum;

[0012] Determine the first sample serial numbers corresponding to the same second sample serial number for each of the minimum distance sums as the respective first sample serial numbers corresponding to the second sample serial number.

[0013] In one embodiment, after the step of respectively sending the respective first sample serial numbers corresponding to each second sample serial number to each client, the method further includes:

[0014] Obtain the sample clustering parameters of the server;

[0015] When it is determined according to the sample clustering parameters that the samples do not meet the iteration condition, send a first message to each client. When the client receives the first message, the client determines the corresponding ordinary sample and the central sample of the second sample serial number as samples in the same cluster.

[0016] In one embodiment, the step of obtaining the sample clustering parameters of the server includes:

[0017] When it is determined according to the sample clustering parameters that the samples meet the iteration condition, send a second message to each client. When the client receives the second message, update the ordinary samples of the respective first sample serial numbers corresponding to the second sample serial number and the central sample corresponding to the second sample serial number as central samples, and send the updated distance and the sample serial number corresponding to the updated distance to the server. The updated distance is the distance between the ordinary sample and the updated central sample;

[0018] Return to execute the step of the server receiving the sample serial numbers of the samples in the data set sent by each client and the distance corresponding to the sample serial number.

[0019] In one embodiment, the sample clustering parameters include at least one of the number of times of clustering samples by the server and the sum of squared errors of sample clustering, and the iteration condition includes at least one of the number of clustering times reaching a preset number of times and the sum of squared errors being less than a preset threshold.

[0020] To achieve the above object, the present invention further provides a method for clustering samples, and the method for clustering samples includes:

[0021] The client determines the distance between the ordinary sample and the central sample in the data set, and determines the first sample serial number of the ordinary sample corresponding to the distance and the second sample serial number of the central sample;

[0022] Send the sample serial numbers and the distances corresponding to the sample serial numbers to the server, where the sample serial numbers include the first sample serial number and the second sample serial number, and the server determines each first sample serial number corresponding to each second sample serial number according to the sample serial numbers and the distances corresponding to the sample serial numbers sent by each client;

[0023] Receive each first sample serial number corresponding to each second sample serial number fed back by the server;

[0024] Determine the ordinary samples and the central samples corresponding to the second sample serial numbers as the same cluster samples.

[0025] In one embodiment, after the step of receiving each first sample serial number corresponding to each second sample serial number fed back by the server, the method further includes:

[0026] Receive the information fed back by the server;

[0027] When the information is the first information, execute the step of determining the ordinary samples and the central samples corresponding to the second sample serial numbers as the same cluster samples, where when it is determined according to the sample clustering parameters of the server that the samples do not meet the iteration condition, send the first information to the client.

[0028] In one embodiment, after the step of receiving the information fed back by the server, the method further includes:

[0029] When the information is the second information, update the ordinary samples of each first sample serial number corresponding to the second sample serial number and the central sample corresponding to the second sample serial number to the central samples, where when it is determined according to the sample clustering parameters of the server that the samples meet the iteration condition, send the second information to the client;

[0030] Determine the distance between the ordinary sample and the updated central sample to obtain the updated distance;

[0031] Determine the first sample serial number and the third sample serial number corresponding to the updated distance, and send the updated distance and the first sample serial number and the third sample serial number corresponding to the updated distance to the server.

[0032] In one embodiment, the step of the client determining the distance between the ordinary sample and the central sample in the data set includes:

[0033] Assign values to each ordinary sample in the data set to obtain the mean vector corresponding to each ordinary sample, and determine the central sample among each ordinary sample;

[0034] Determine the distance between the ordinary sample and the central sample according to the mean vector of the ordinary sample and the mean vector of the central sample.

[0035] In one embodiment, the step of determining the distance between the ordinary sample and the central sample in the data set includes:

[0036] Obtain a random number and determine the true distance between the ordinary sample and the central sample in the data set;

[0037] Determine the distance between the ordinary sample and the central sample according to the random number and the true distance.

[0038] To achieve the above object, the present invention further provides a server, which includes:

[0039] A receiving module, configured to receive the sample numbers of the samples in the data set sent by each client and the distance corresponding to the sample numbers, where the distance is the distance between the ordinary sample corresponding to the sample number and the central sample in the data set;

[0040] A determining module, configured to determine each first sample number corresponding to each second sample number according to each of the distances, where the sample numbers include the first sample number of the ordinary sample and the second sample number of the central sample, and the data set includes multiple central samples;

[0041] A sending module, configured to send each of the first sample numbers corresponding to each second sample number to each of the clients respectively, where the client determines the ordinary sample and the central sample corresponding to the second sample number as samples in the same cluster.

[0042] To achieve the above object, the present invention further provides a client, which includes:

[0043] A determining module, configured to determine the distance between the ordinary sample and the central sample in the data set, and determine the first sample number of the ordinary sample corresponding to the distance and the second sample number of the central sample;

[0044] A sending module, configured to send the sample number and the distance corresponding to the sample number to the server, where the sample numbers include the first sample number and the second sample number, and the server determines each first sample number corresponding to each second sample number according to the sample numbers and the distances corresponding to the sample numbers sent by each client;

[0045] A receiving module, configured to receive each of the first sample numbers corresponding to each second sample number fed back by the server;

[0046] The determining module is configured to determine the general sample and the central sample corresponding to the second sample number as the samples in the same cluster.

[0047] To achieve the above object, the present invention further provides a device, which includes a memory, a processor, and a clustering program stored in the memory and executable on the processor. When the clustering program is executed by the processor, the steps of the sample clustering method described above are implemented, or the steps of the sample clustering method described above are executed.

[0048] To achieve the above object, the present invention further provides a readable storage medium storing a clustering program, and when the clustering program is executed by a processor, the steps of the sample clustering method described above are implemented.

[0049] The present invention provides a sample clustering method, a server, a client, a device, and a readable storage medium. The server receives the sample numbers of the samples in the data set sent by each client and the distances corresponding to the sample numbers, and determines each first sample number corresponding to each second sample number according to each distance, and then sends each first sample number corresponding to each second sample number to each client respectively, so that the client determines the central sample and the general sample corresponding to the second sample number as the samples in the same cluster. The client of the present invention sends the second sample number of the central sample in the data set, the first sample number of the general sample, and the distance between the general sample and the central sample in the data set to the server, so that the server determines each general sample in the data set that is of the same category as the central sample according to the sample number and the distance. Compared with the prior art where data leakage occurs when the server clusters samples pair by pair, the server of the present invention can cluster each sample in the data set without obtaining the real sample data of the client, avoiding data leakage of the client. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 It is a schematic hardware structure diagram of the server / client / device of the hardware operating environment related to the solution of the embodiment of the present invention;

[0051] Figure 2 It is a schematic flowchart of the first embodiment of the sample clustering method of the present invention;

[0052] Figure 3 For Figure 2 It is a detailed flowchart of step S20 in

[0053] Figure 4 It is a schematic flowchart of the second embodiment of the sample clustering method of the present invention;

[0054] Figure 5 It is a schematic flowchart of the third embodiment of the sample clustering method of the present invention;

[0055] Figure 6 Schematic diagram of the functional modules of the server side of the present invention;

[0056] Figure 7 Schematic diagram of the functional modules of the client side of the present invention.

[0057] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Specific embodiments

[0058] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0059] Refer to Figure 1 , Figure 1 The hardware structure diagram of the hardware operating environment of the server / client / terminal involved in the embodiment solution of the present invention.

[0060] As Figure 1 shown, the server / client / device may include: a processor 1001, such as a CPU, a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a Wi-Fi interface). The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0061] Those skilled in the art can understand that Figure 1 the structure shown in

[0062] does not constitute a limitation on the server, client or device, and may include more or fewer components than shown in the figure, or combine some components, or arrange different components. Figure 1 As

[0063] shown in Figure 1In the server or service end shown, the network interface 1004 is mainly used to connect to the background service end and communicate data with the background service end; the user interface 1003 is mainly used to connect to the client and communicate data with the client; and the processor 1001 can be used to call the clustering program stored in the memory 1005 and perform the following operations:

[0064] The server receives the sample numbers of the samples in the data set sent by each client and the distances corresponding to the sample numbers, where the distance is the distance between the normal sample corresponding to the sample number and the central sample in the data set;

[0065] Determine each first sample number corresponding to each second sample number according to each of the distances, where the sample numbers include the first sample number of the normal sample and the second sample number of the central sample, and the data set includes multiple central samples;

[0066] Send each first sample number corresponding to each second sample number to each of the clients, where the client determines the normal sample and the central sample corresponding to the second sample number as co-cluster samples.

[0067] In an embodiment, the processor 1001 can call the clustering program stored in the memory 1005 and also perform the following operations:

[0068] Determine a target distance among each of the distances, and determine the sum of the distances of each of the target distances, where the first sample number and the second sample number corresponding to each of the target distances are the same;

[0069] Determine the minimum sum of distances among the sums of the distances corresponding to the same first sample numbers;

[0070] Determine each first sample number corresponding to each second sample number as the first sample numbers corresponding to each minimum sum of distances of the same second sample number.

[0071] In an embodiment, the processor 1001 can call the clustering program stored in the memory 1005 and also perform the following operations:

[0072] Obtain the sample clustering parameters of the server;

[0073] When it is determined according to the sample clustering parameters that the samples do not meet the iteration condition, send a first message to each of the clients, where when the client receives the first message, the client determines the normal sample and the central sample corresponding to the second sample number as co-cluster samples.

[0074] In one embodiment, the processor 1001 may call the clustering program stored in the memory 1005 and further perform the following operations:

[0075] When it is determined according to the sample clustering parameter that the sample meets the iteration condition, send second information to each of the clients. When the client receives the second information, update the normal samples corresponding to each first sample serial number and the central sample corresponding to the second sample serial number to central samples, and send the updated distance and the sample serial number corresponding to the updated distance to the server, where the updated distance is the distance between the normal sample and the updated central sample;

[0076] Return to execute the step of the server receiving the sample serial numbers of the samples in the data set sent by each client and the distances corresponding to the sample serial numbers.

[0077] In one embodiment, the processor 1001 may call the clustering program stored in the memory 1005 and further perform the following operations:

[0078] The sample clustering parameter includes at least one of the number of times the server clusters the samples and the sum of squared errors of the sample clustering, and the iteration condition includes at least one of the number of clustering times reaching a preset number and the sum of squared errors being less than a preset threshold.

[0079] In one embodiment, the processor 1001 may call the clustering program stored in the memory 1005 and further perform the following operations:

[0080] The client determines the distance between the normal sample and the central sample in the data set, and determines the first sample serial number of the normal sample corresponding to the distance and the second sample serial number of the central sample;

[0081] Send the sample serial number and the distance corresponding to the sample serial number to the server, where the sample serial number includes the first sample serial number and the second sample serial number, and the server determines each first sample serial number corresponding to each second sample serial number according to the sample serial numbers and the distances corresponding to the sample serial numbers sent by each client;

[0082] Receive each first sample serial number corresponding to each second sample serial number fed back by the server;

[0083] Determine the normal sample and the central sample corresponding to the second sample serial number as the same cluster samples.

[0084] In one embodiment, the processor 1001 may call the clustering program stored in the memory 1005 and further perform the following operations:

[0085] Receive the information fed back by the server;

[0086] When the information is the first information, perform the step of determining the normal sample and the central sample corresponding to the second sample serial number as the samples in the same cluster, wherein when it is determined according to the sample clustering parameter of the server that the sample does not meet the iteration condition, send the first information to the client.

[0087] In one embodiment, the processor 1001 may call the clustering program stored in the memory 1005 and further perform the following operations:

[0088] When the information is the second information, update the normal samples corresponding to each first sample serial number of the second sample serial number and the central sample corresponding to the second sample serial number as the central samples, wherein when it is determined according to the sample clustering parameter of the server that the sample meets the iteration condition, send the second information to the client;

[0089] Determine the updated distance by determining the distance between the normal sample and the updated central sample;

[0090] Determine the first sample serial number and the third sample serial number corresponding to the updated distance, and send the updated distance and the first sample serial number and the third sample serial number corresponding to the updated distance to the server.

[0091] In one embodiment, the processor 1001 may call the clustering program stored in the memory 1005 and further perform the following operations:

[0092] Assign values to each of the normal samples in the dataset to obtain the mean vector corresponding to each of the normal samples, and determine the central sample among each of the normal samples;

[0093] Determine the distance between the normal sample and the central sample according to the mean vector of the normal sample and the mean vector of the central sample.

[0094] In one embodiment, the processor 1001 may call the clustering program stored in the memory 1005 and further perform the following operations:

[0095] Obtain a random number and determine the true distance between the normal sample and the central sample in the dataset;

[0096] Determine the distance between the normal sample and the central sample according to the random number and the true distance.

[0097] Based on the above hardware structure of the server, various embodiments of the clustering method of the samples of the present invention are proposed.

[0098] The present invention provides a clustering method for samples.

[0099] Refer to Figure 2 , Figure 2 This is the first embodiment of the clustering method for the samples of the present invention. The clustering method for the samples includes:

[0100] Step S10, the server receives the sample numbers of the samples in the data set sent by each client and the distances corresponding to the sample numbers, where the distance is the distance between the ordinary sample corresponding to the sample number and the central sample in the data set;

[0101] In this embodiment, the client includes multiple samples, and the samples can be obtained by the server aligning the data provided by each client. Data alignment is the data alignment method in vertical federated learning, that is, after the server receives the feature vectors representing the data sent by each client, it first determines the distances between the feature vectors, and determines the feature vectors with distances less than the preset distance as the feature vectors belonging to the same class. The feature vector represents a numerical value. The server then performs an intersection operation on the numerical values of the feature vectors in the same class (the intersection operation of the feature vectors in the same class is to perform data alignment on the same class of data), and feeds back the feature vectors after the intersection operation to each client, so that the client can obtain the samples according to the feature vectors after the intersection operation and configure sample numbers for the samples.

[0102] The client can divide the samples in the data set into K groups, and set one sample center in a group of samples, that is, set K sample centers for K groups of samples. Each client sets K groups, so the number of sample centers set by each client is the same. K is an integer and K is greater than or equal to 2. The samples in the data set can be defined as ordinary samples, and the client can randomly determine the central samples among the ordinary samples. Therefore, the samples in the data set may have two identities. One identity is as the central sample of a certain group, and the other identity is as the ordinary sample of other groups. The client will set different sample numbers for each sample in the data set. To distinguish the samples with two identities, the client first configures the first sample number for each ordinary sample, and then configures the second sample number for the ordinary sample that is the central sample. It can be understood that the sample number of the ordinary sample is the first sample number, and the sample number of the central sample is the second sample number.

[0103] After the client determines the central samples, it first assigns values to each ordinary sample to obtain the mean vector corresponding to the ordinary sample. The client calculates the distance between the central sample and each ordinary sample through the mean vector of the sample. In fact, the distance corresponds to a first sample number and a second sample number. The client associates the first sample number, the second sample number, and the distance, and thus sends the association information carrying the distance to the client.

[0104] Further, to avoid the disclosure of the real distance, the client can generate a random number. After obtaining the real distance, the client determines the distance to be sent to the server based on the random number and the real distance, that is, the real distance is superimposed with the random number to obtain the distance to be sent to the client.

[0105] The server receives the sample numbers of the samples in the data set sent by each client and the distances corresponding to the sample numbers.

[0106] Step S20: Determine each first sample number corresponding to each second sample number according to each of the distances. The sample numbers include the first sample numbers of ordinary samples and the second sample numbers of central samples, and the data set includes multiple central samples.

[0107] After the server obtains each distance, since each distance corresponds to a first sample number and a second sample number, the server can determine each first sample number corresponding to each second sample number according to the distance, the first sample number corresponding to the distance, and the second sample number. Specifically, referring to Figure 3 , that is, step S20 includes:

[0108] Step S21: Determine the target distances among each of the distances, and determine the sum of the distances of each of the target distances. The first sample numbers and the second sample numbers corresponding to each of the target distances are the same.

[0109] Step S22: Determine the minimum sum of distances among the sums of distances corresponding to the same first sample numbers.

[0110] Step S23: Determine the first sample numbers corresponding to each of the minimum sums of distances with the same second sample number as each of the first sample numbers corresponding to the second sample number.

[0111] Each client has the same first sample serial number for the sample configuration. For example, the serial numbers configured by Client A are 1 - K, and the serial numbers configured by Client B are also 1 - K. And each client has the same second sample serial number for the configuration of the central sample. For example, the second sample serial numbers configured by each client are "Central 1", "Central 2". Therefore, the server can determine the target distance among various distances. The first sample serial numbers and the second sample serial numbers corresponding to each target distance are the same. The server then sums up the target distances to obtain the distance sum. The target distances with the same first sample serial number and the same second sample serial number are grouped together. The server sums up the target distances within this group to obtain the distance sum. Since there are multiple central samples in the dataset, there are multiple distance sums corresponding to one sample serial number, that is, there are multiple groups of target distances. The server determines the minimum distance sum from the various distance sums corresponding to the first sample serial number. The minimum distance sum corresponds to a second sample serial number. From this, the second sample serial number closest to each first sample serial number can be determined, that is, the server determines the central sample closest to the ordinary sample. The server then counts the central samples closest to the ordinary samples to obtain each ordinary sample closest to the central sample. Since the server cannot know the sample data, the server determines each first sample serial number corresponding to each second sample serial number. The following is an example to illustrate in detail the determination of each first sample serial number corresponding to the second sample serial number.

[0112] There are three samples A, B, and C in the dataset of Client1 (Client 1). A, B, and C are the first sample serial numbers. Among them, the samples corresponding to B and C are used as the central samples, and the second sample serial numbers are "Central 1" and "Central 2". There are three samples A, B, and C in the dataset of Client2 (Client 2). A, B, and C are the first sample serial numbers. Among them, the samples corresponding to B and C are used as the central samples, and the second sample serial numbers are "Central 1" and "Central 2".

[0113] In the dataset of Client1, the distance of "A - Central 1" is 1, the distance of "B - Central 1" is 0, and the distance of "C - Central 1" is 3; the distance of "A - Central 2" is 3, the distance of "B - Central 1" is 2, and the distance of "C - Central 1" is 0. In the dataset of Client2, the distance of "A - Central 2" is 1, the distance of "B - Central 1" is 0, and the distance of "C - Central 1" is 7; the distance of "A - Central 2" is 5, the distance of "B - Central 1" is 7, and the distance of "C - Central 1" is 0.

[0114] The data sent by Client1 to the server is "A - Central 1: 1 + r1", "B - Central 1: 0 + r1", "C - Central 1: 3 + r1", "A - Central 2: 3 + r1", "B - Central 2: 2 + r1", "C - Central 2: 0 + r1". Among them, r1 is a random number.

[0115] The data sent by Client2 to the server is "A - Zhong1: 2 + r2", "B - Zhong1: 0 + r2", "C - Zhong1: 7 + r2", "A - Zhong2: 5 + r2", "B - Zhong2: 7 + r2", "C - Zhong2: 0 + r2". Among them, r2 is a random number.

[0116] The server calculates the respective target distances corresponding to the same first serial number and the same second serial number, that is, the distances corresponding to A - Zhong1 are "1 + r1" and "2 + r2", the distances corresponding to B - Zhong1 are "0 + r1" and "0 + r2", the distances corresponding to C - Zhong1 are "3 + r1" and "7 + r2", the distances corresponding to A - Zhong2 are "3 + r1" and "5 + r2", the distances corresponding to B - Zhong2 are "2 + r1" and "7 + r2", and the distances corresponding to C - Zhong1 are "0 + r1" and "0 + r2".

[0117] The server determines the sum of the distances corresponding to the first sample serial number, that is, the sum of the distances corresponding to the first sample serial number A is 3 + r1 + r2 and 8 + r1 + r2; the sum of the distances corresponding to B is 0 + r1 + r2 and 9 + r1 + r2; the sum of the distances corresponding to C is 9 + r1 + r2 and 0 + r1 + r2.

[0118] The server then determines the minimum sum of distances corresponding to each first sample serial number. For example, the minimum sum of distances corresponding to A is 3 + r1 + r2 (3 + r1 + r2 is less than 8 + r1 + r2), the minimum sum of distances corresponding to B is 0 + r1 + r2 (0 + r1 + r2 is less than 9 + r1 + r2), and the minimum sum of distances corresponding to C is 0 + r1 + r2 (0 + r1 + r2 is less than 9 + r1 + r2). Therefore, Zhong1 (Zhong1 is the second sample serial number) corresponds to A and B (A and B are the first sample serial numbers), and Zhong2 (Zhong2 is the second sample serial number) corresponds to C.

[0119] Step S30: Send each of the first sample serial numbers corresponding to each second sample serial number to each client, where the client determines the normal sample and the central sample corresponding to the second sample serial number as the same cluster samples.

[0120] After the server determines each first sample serial number corresponding to each second sample serial number, it sends each first sample serial number corresponding to each second sample serial number to the client, so that the client determines the normal sample and the central sample corresponding to the second sample serial number as the same cluster samples. Each sample belonging to the same cluster samples is a type of sample, and the client can complete the clustering of the samples. Still taking the above example to illustrate the client's determination of the same cluster samples. It should be noted that the normal sample corresponding to the second sample serial number is the normal sample of the first sample serial number corresponding to the second sample serial number.

[0121] For example, if Clent1 receives (Zhong1, A, B) and (Zhong2, C) sent by the server, then Clent1 clusters the samples corresponding to A and B into one center, that is, regards the two samples A and B as one type of sample, and regards C as one type of sample.

[0122] In the technical solution provided in this embodiment, the server receives the sample numbers of the samples in the data set sent by each client and the distances corresponding to the sample numbers, determines each first sample number corresponding to each second sample number according to each distance, and then sends each first sample number corresponding to each second sample number to each client respectively, so that the client determines the central sample corresponding to the second sample number and the normal samples as the same cluster samples. In the present invention, the client sends the second sample number of the central sample in the data set, the first sample number of the normal sample, and the distance between the normal sample and the central sample in the data set to the server, so that the server determines each normal sample in the data set that is of the same type as the central sample according to the sample number and the distance. Compared with the prior art where there is data leakage when the server clusters sample by sample, in the present invention, the server can cluster each sample in the data set without obtaining the real sample data of the client, avoiding data leakage of the client.

[0123] Refer to Figure 4 , Figure 4 This is the second embodiment of the clustering method of the samples of the present invention. Based on the first embodiment, after the step 30, it further includes:

[0124] Step S40, obtaining the sample clustering parameters of the server;

[0125] Step S50, when it is determined according to the sample clustering parameters that the samples do not meet the iteration condition, sending a first message to each of the clients. When the client receives the first message, the client determines the normal samples corresponding to each first sample number of the second sample number and the central sample corresponding to the second sample number as the same cluster samples;

[0126] Step S60, when it is determined according to the sample clustering parameters that the samples meet the iteration condition, sending a second message to each of the clients. When the client receives the second message, it updates the normal samples corresponding to each first sample number of the second sample number and the central sample corresponding to the second sample number to the central samples, and sends the updated distance and the sample number corresponding to the updated distance to the server. The updated distance is the distance between the normal sample and the updated central sample;

[0127] In this embodiment, the server needs to determine whether the client needs to resend the distance for iterative clustering of samples according to the sample clustering parameters. The sample clustering parameters include at least one of the number of times the server clusters the samples and the sum of squared errors of sample clustering. The server determines the number of times of each first sample serial number corresponding to each second sample serial number according to the sample serial numbers provided by the client and the distances corresponding to the sample serial numbers, which is the number of clustering times. The number of times of sample clustering refers to the number of times the server clusters the sample serial numbers and distances sent by the client. The calculation method of the sum of squared errors of the sample is as follows: Determine the difference between the distance corresponding to the current central sample and the distance corresponding to the previous central sample, the square of a difference corresponding to a central sample, and the sum of the squares of all the differences is the sum of squared errors. The fitted value can be the fitted value of the clustering model, and the clustering model can replace the server to perform the sample clustering operation. Specifically, the current number of times of sample clustering performed by the server may be the third time, and the preset number of times is five times. The server needs to cluster the samples of the client again. The client will update the central sample corresponding to the second sample serial number and the ordinary samples of each first sample serial number corresponding to the second sample serial number to the central sample, and the central sample vector will also be updated accordingly. For example, if AB is updated to the central sample and C is also the central sample, the mean vector of the updated central samples is (A1 + B1) / 2, where A1 is the mean vector of the sample corresponding to A and B1 is the mean vector of the sample corresponding to B. Then the sum of the mean vectors is ((A1 + B1) / 2) + C1; the sum of the previous mean vectors is A1 + B1 + C1. Therefore, the change in the sum of the mean vectors is (A1 + B1) / 2, (A1 + B1) 2 / 4 is the sum of squared errors.

[0128] After obtaining the sample clustering parameters, the server determines whether the samples meet the iteration conditions according to the sample clustering parameters. The iteration conditions include at least one of the number of clustering times reaching the preset number of times and the sum of squared errors being less than the preset threshold.

[0129] If the samples do not meet the iteration conditions, the server sends the first information to the client, so that when the client receives the first information, it determines the ordinary samples of each first sample serial number corresponding to the second sample serial number and the central sample corresponding to the second sample serial number as samples in the same cluster, so as to complete the clustering of the samples in the dataset.

[0130] When the sample meets the iteration condition, the server sends the second information to the client, so that when the client receives the first information, it first updates the central sample corresponding to the second sample serial number and the ordinary samples corresponding to each first sample serial number of the second sample to the central sample, and then calculates the distance between each ordinary sample and the central sample to obtain the updated distance. The client then sends the updated distance and the sample signal corresponding to the updated distance to the server, so that the server performs sample clustering again. It should be noted that the first information refers to the information carrying the message that the sample clustering is not completed, and the second information refers to the information carrying the message that the sample clustering is completed.

[0131] Taking the dataset of clent1 including three samples A, B, and C as an example, if the updated central sample is (A + B) / 2, then the third sample serial number of the updated central sample is (A + B) / 2, and the middle one is C. Then the distance between (A / 2 + B / 2) and the first central sample is 0, the distance between C and the first central sample is 2.5, the distance between (A / 2 + B / 2) and the second central sample is 2.5, and the distance between C and the first central sample is 0. The client sends ((A / 2 + B / 2), the first central sample, 0 + r1), (C, the first central sample, 2.5 + r1), ((A / 2 + B / 2), the second central sample, 2.5 + r1), (C, the second central sample, 0 + r1) to the server. The server then returns to execute the step of receiving the sample serial numbers in the dataset sent by each client and the distances corresponding to the sample serial numbers. That is, the server executes steps S10 - S40 again.

[0132] In the technical solution provided in this embodiment, the server obtains the sample clustering parameters, and determines whether the sample meets the iteration condition through the sample clustering parameters, so that the server accurately clusters the samples of the client.

[0133] The present invention also provides a method for clustering samples.

[0134] Refer to Figure 5 , Figure 5 This is the third embodiment of the sample clustering method of the present invention. The sample clustering method includes:

[0135] Step S100, the client determines the distance between the ordinary sample and the central sample in the dataset, and determines the first sample serial number of the ordinary sample corresponding to the distance and the second sample serial number of the central sample;

[0136] Step S110, sending the sample serial number and the distance corresponding to the sample serial number to the server, where the sample serial number includes the first sample serial number and the second sample serial number, and the server determines each first sample serial number corresponding to each second sample serial number according to the sample serial numbers and the distances corresponding to the sample serial numbers sent by each client;

[0137] Step S120: Receive each of the first sample numbers corresponding to the second sample numbers fed back by the server.

[0138] Step S130: Determine the normal samples and the central samples corresponding to the second sample numbers as samples in the same cluster.

[0139] In this embodiment, the executing entity is a client. The client is provided with a data set, which includes multiple samples, and each sample can be defined as a normal sample. The client can determine the central sample among the normal samples. The client configures corresponding sample numbers for each sample, calculates the distance between each normal sample and the central sample, and then sends the distance and the corresponding sample number to the server, so that the server clusters the samples in the data set according to the sample numbers provided by each client and the distances corresponding to the sample numbers. That is, the server determines each of the first sample numbers corresponding to each second sample number, and then sends each of the first sample numbers corresponding to each second sample number to the client. The client determines the normal samples corresponding to the first sample numbers of the second sample number and the central sample corresponding to the second sample number as samples in the same cluster to complete the clustering of the samples.

[0140] The samples in the data set are all samples obtained after data alignment. The client first converts the data into feature vectors, and then sends the feature vectors to the server. The server performs an intersection operation on the same type of feature vectors, and then feeds back the feature vectors after the intersection operation to the client. The client then determines the samples according to the feature vectors after the intersection operation.

[0141] In this embodiment, the process of the client determining the distance between the normal sample and the central sample, the server determining each of the first sample numbers corresponding to each second sample number according to the distance, and the client clustering the samples corresponding to each of the first sample numbers of the second sample number refers to the above embodiment and will not be elaborated here.

[0142] In one embodiment, after step S120, it further includes:

[0143] C1: Receive the information fed back by the server;

[0144] C2: When the information is the first information, execute the step of determining the normal samples corresponding to each of the first sample numbers of the second sample number and the central sample corresponding to the second sample number as samples in the same cluster. Wherein, when it is determined according to the sample clustering parameters of the server that the samples do not meet the iteration condition, send the first information to the client.

[0145] C3. When the information is the second information, update the normal samples corresponding to each first sample serial number and the central sample corresponding to the second sample serial number to central samples, where when it is determined according to the sample clustering parameters of the server that the samples meet the iteration condition, send the second information to the client;

[0146] C4. Determine the updated distance by obtaining the distance between the normal sample and the updated central sample;

[0147] C5. Determine the first sample serial number and the third sample serial number corresponding to the updated distance, and send the updated distance and the first sample serial number and the third sample serial number corresponding to the updated distance to the server.

[0148] In this embodiment, the first information, the second information, the iteration condition, the server sending the first information, the server sending the second information, the operations after the client receives the first information, and the operations after the client receives the second information specifically refer to the descriptions in the above embodiments, and will not be elaborated here.

[0149] In one embodiment, step S100 includes:

[0150] Assign a value to each normal sample in the dataset to obtain the mean vector corresponding to each normal sample, and determine the central sample among the normal samples;

[0151] According to the mean vector of the normal sample and the mean vector of the central sample, determine the distance between the normal sample and the central sample.

[0152] In this embodiment, the determination of the distance between the normal sample and the central sample refers to the description in the above embodiment, and will not be elaborated here.

[0153] In one embodiment, step S100 includes:

[0154] Obtain a random number, and determine the true distance between the normal sample and the central sample in the dataset;

[0155] According to the random number and the true distance, determine the distance between the normal sample and the central sample.

[0156] In this embodiment, the distance to be determined between the normal sample and the central sample is the true distance between the two, and the distance obtained by superimposing the true distance and the random number is the distance sent by the client to the server. The determination of the true distance, and the determination of the random number and the true distance refer to the descriptions in the above embodiments, and will not be elaborated here.

[0157] The present invention also provides a server.

[0158] Reference Figure 6 , Figure 6 which is a schematic diagram of the functional modules of the server of the present invention.

[0159] As Figure 6 shown, the server includes:

[0160] A receiving module 10, configured to receive the sample numbers of the samples in the data set sent by each client and the distances corresponding to the sample numbers, where the distance is the distance between the normal sample corresponding to the sample number and the central sample in the data set;

[0161] A determining module 20, configured to determine, according to each of the distances, each first sample number corresponding to each second sample number, where the sample numbers include the first sample numbers of the normal samples and the second sample numbers of the central samples, and the data set includes a plurality of central samples;

[0162] A sending module 30, configured to send each of the first sample numbers corresponding to each second sample number to each of the clients respectively, where the client determines the normal sample corresponding to the second sample number and the central sample as co-cluster samples.

[0163] In an embodiment, the server further includes:

[0164] A determining module 20, configured to determine a target distance among each of the distances and determine the sum of the distances of each of the target distances, where the first sample numbers and the second sample numbers corresponding to each of the target distances are the same;

[0165] A determining module 20, configured to determine the minimum sum of distances among the sums of the distances corresponding to the same first sample numbers;

[0166] A determining module 20, configured to determine, as each of the first sample numbers corresponding to each second sample number, the first sample numbers corresponding to each of the minimum sums of distances of the same second sample numbers.

[0167] In an embodiment, the server further includes:

[0168] An obtaining module, configured to obtain the sample clustering parameters of the server;

[0169] A sending module 30, configured to send a first message to each of the clients when it is determined according to the sample clustering parameters that the samples do not meet the iteration condition, where when the client receives the first message, the client determines the normal sample corresponding to the second sample number and the central sample as co-cluster samples.

[0170] In an embodiment, the server further includes:

[0171] A sending module 30, configured to send second information to each of the clients when it is determined according to the sample clustering parameter that the sample meets the iteration condition. When each client receives the second information, it updates the normal samples corresponding to each first sample number and the central sample corresponding to the second sample number to central samples, and sends the updated distance and the sample number corresponding to the updated distance to the server, where the updated distance is the distance between the normal sample and the updated central sample;

[0172] An execution module, configured to return and execute the step of the server receiving the sample numbers in the data set sent by each client and the distance corresponding to the sample numbers.

[0173] Wherein, the function implementation of each module in the above server corresponds to each step in the embodiment of the sample clustering method, and its function and implementation process will not be elaborated here one by one.

[0174] The present invention further provides a client.

[0175] Referring to Figure 7 , Figure 7 which is a schematic diagram of the function modules of the client of the present invention.

[0176] As Figure 7 shown, the client includes:

[0177] A determination module 10, configured to determine the distance between a normal sample and a central sample in a data set, and determine the first sample number of the normal sample corresponding to the distance and the second sample number of the central sample;

[0178] A sending module 20, configured to send the sample number and the distance corresponding to the sample number to a server, where the sample number includes the first sample number and the second sample number, and the server determines each first sample number corresponding to each second sample number according to the sample number and the distance corresponding to the sample number sent by each client;

[0179] A receiving module 30, configured to receive each first sample number corresponding to each second sample number fed back by the server;

[0180] A determination module 10, configured to determine the normal sample and the central sample corresponding to the second sample number as co-cluster samples.

[0181] In an embodiment, the client further includes:

[0182] A receiving module 30, configured to receive the information fed back by the server;

[0183] An execution module, configured to, when the information is the first information, execute the step of determining the normal sample corresponding to the second sample serial number and the central sample as the same cluster samples, wherein when it is determined according to the sample clustering parameters of the server that the samples do not meet the iteration condition, send the first information to the client.

[0184] In one embodiment, the client further includes:

[0185] An update module, configured to, when the information is the second information, update the normal samples of each first sample serial number corresponding to the second sample serial number and the central sample corresponding to the second sample serial number as central samples, wherein when it is determined according to the sample clustering parameters of the server that the samples meet the iteration condition, send the second information to the client;

[0186] A determination module 10, configured to determine an updated distance of the distance between the normal sample and the updated central sample;

[0187] A determination module 10, configured to determine the first sample serial number and the third sample serial number corresponding to the updated distance, and send the updated distance and the first sample serial number and the third sample serial number corresponding to the updated distance to the server.

[0188] In one embodiment, the client further includes:

[0189] An assignment module, configured to assign each of the normal samples in the dataset to obtain a mean vector corresponding to each of the normal samples, and determine a central sample among the normal samples;

[0190] A determination module 10, configured to determine the distance between the normal sample and the central sample according to the mean vector of the normal sample and the mean vector of the central sample.

[0191] In one embodiment, the client further includes:

[0192] An acquisition module 20, configured to acquire a random number and determine a true distance between the normal sample and the central sample in the dataset;

[0193] A determination module 10, configured to determine the distance between the normal sample and the central sample according to the random number and the true distance.

[0194] Wherein, the function implementation of each module in the above client corresponds to each step in the embodiment of the sample clustering method above, and its function and implementation process will not be elaborated here one by one.

[0195] The present invention also provides a readable storage medium, on which a clustering program is stored. When the clustering program is executed by a processor, the steps of the clustering method of the sample as described in any of the above embodiments are implemented.

[0196] The specific embodiments of the medium of the present invention are basically the same as those of the above embodiments of the clustering method of the sample, and will not be described in detail here.

[0197] It should be noted that in this article, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or system including the element.

[0198] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments.

[0199] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal device (which may be a mobile phone, a computer, a server, an air conditioner, or a network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0200] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. A clustering method for samples, characterized in that, the clustering method for samples includes: The server receives the sample numbers of the samples in the data set sent by each client and the distances corresponding to the sample numbers, where the distance is the distance between the ordinary sample corresponding to the sample number and the central sample in the data set; Determine each first sample number corresponding to each second sample number according to each of the distances, where the sample numbers include the first sample numbers of ordinary samples and the second sample numbers of central samples, and the data set includes multiple central samples; Send each first sample number corresponding to each second sample number to each of the clients, where the client determines the ordinary sample corresponding to the second sample number and the central sample as samples in the same cluster; The step of determining each first sample number corresponding to each second sample number according to each of the distances includes: Determine a target distance among each of the distances, and determine the sum of the distances of each of the target distances, where the first sample numbers and the second sample numbers corresponding to each of the target distances are the same; Determine the minimum sum of distances among the sums of distances corresponding to the same first sample numbers; Determine the first sample numbers corresponding to each of the minimum sums of distances of the same second sample numbers as each of the first sample numbers corresponding to the second sample numbers.

2. The clustering method for samples according to claim 1, characterized in that, after the step of sending each first sample number corresponding to each second sample number to each of the clients, it further includes: Obtain the sample clustering parameters of the server; When it is determined according to the sample clustering parameters that the samples do not meet the iteration condition, send a first message to each of the clients, where when the client receives the first message, the client determines the ordinary sample corresponding to the second sample number and the central sample as samples in the same cluster.

3. The clustering method for samples according to claim 2, characterized in that, the step of obtaining the sample clustering parameters of the server includes: When it is determined according to the sample clustering parameters that the samples meet the iteration condition, send a second message to each of the clients, where when the client receives the second message, update the ordinary samples of each first sample number corresponding to the second sample number and the central sample corresponding to the second sample number to central samples, and send the updated distance and the sample number corresponding to the updated distance to the server, where the updated distance is the distance between the ordinary sample and the updated central sample; Return to execute the step of the server receiving the sample numbers of the samples in the data set sent by each client and the distances corresponding to the sample numbers.

4. The clustering method for samples according to claim 2 or 3, characterized in that, the sample clustering parameters include at least one of the number of times the server clusters the samples and the sum of squared errors of sample clustering, and the iteration condition includes at least one of the number of clustering times reaching a preset number and the sum of squared errors being less than a preset threshold.

5. A clustering method for samples, It is characterized in that the clustering method of the samples includes: the client determines the distance between the ordinary samples and the central samples in the data set, and determines the first sample numbers of the ordinary samples corresponding to the distances and the second sample numbers of the central samples; sending the sample numbers and the distances corresponding to the sample numbers to the server, wherein the sample numbers include the first sample numbers and the second sample numbers, the server determines the target distance among the distances, and determines the sum of the distances of each target distance, and the first sample numbers and the second sample numbers corresponding to each target distance are the same; among the sums of the distances corresponding to the same first sample numbers, the minimum sum of the distances is determined; the first sample numbers corresponding to each minimum sum of the distances with the same second sample numbers are determined as the first sample numbers corresponding to each second sample number; receiving each first sample number corresponding to the second sample number fed back by the server; determining the ordinary samples and the central samples corresponding to the second sample numbers as the samples in the same cluster.

6. The clustering method of the samples according to claim 5, It is characterized in that after the step of receiving each first sample number corresponding to the second sample number fed back by the server, the method further includes: receiving the information fed back by the server; when the information is the first information, performing the step of determining the ordinary samples and the central samples corresponding to the second sample numbers as the samples in the same cluster, wherein when it is determined according to the sample clustering parameters of the server that the samples do not meet the iteration condition, the first information is sent to the client.

7. The clustering method of the samples according to claim 6, It is characterized in that after the step of receiving the information fed back by the server, the method further includes: when the information is the second information, updating the ordinary samples with the first sample numbers corresponding to the second sample numbers and the central samples corresponding to the second sample numbers as the central samples, wherein when it is determined according to the sample clustering parameters of the server that the samples meet the iteration condition, the second information is sent to the client; determining the distance between the ordinary samples and the updated central samples to obtain the updated distance; determining the first sample number and the third sample number corresponding to the updated distance, and sending the updated distance and the first sample number and the third sample number corresponding to the updated distance to the server.

8. The clustering method of the samples according to claim 5, It is characterized in that the step in which the client determines the distance between the ordinary samples and the central samples in the data set includes: assigning values to each ordinary sample in the data set to obtain the mean vector corresponding to each ordinary sample, and determining the central sample among the ordinary samples; determining the distance between the ordinary sample and the central sample according to the mean vector of the ordinary sample and the mean vector of the central sample.

9. The clustering method of the samples according to claim 5, It is characterized in that the step of determining the distance between the ordinary samples and the central samples in the data set includes: Obtain a random number and determine the true distance between the ordinary sample and the central sample in the dataset; Determine the distance between the ordinary sample and the central sample according to the random number and the true distance.

10. A server, characterized in that, the server includes: a receiving module, configured to receive the sample numbers of the samples in the dataset sent by each client and the distances corresponding to the sample numbers, where the distances are the distances between the ordinary samples corresponding to the sample numbers and the central samples in the dataset; a determining module, configured to determine each first sample number corresponding to each second sample number according to each of the distances, where the sample numbers include the first sample numbers of the ordinary samples and the second sample numbers of the central samples, and the dataset includes multiple central samples; the determining module is further configured to determine a target distance among each of the distances, and determine the sum of the distances of each of the target distances, where the first sample numbers and the second sample numbers corresponding to each of the target distances are the same; determine the minimum sum of distances among the sums of the distances corresponding to the same first sample numbers; and determine each first sample number corresponding to each of the minimum sums of distances of the same second sample numbers as each of the first sample numbers corresponding to the second sample numbers; a sending module, configured to send each of the first sample numbers corresponding to each second sample number to each client respectively, where the client determines the ordinary sample and the central sample corresponding to the second sample number as samples in the same cluster.

11. A client, characterized in that, the client includes: a determining module, configured to determine the distance between the ordinary sample and the central sample in the dataset, and determine the first sample number of the ordinary sample and the second sample number of the central sample corresponding to the distance; a sending module, configured to send the sample number and the distance corresponding to the sample number to the server, where the sample numbers include the first sample number and the second sample number, and the server determines a target distance among each of the distances, and determines the sum of the distances of each of the target distances, where the first sample numbers and the second sample numbers corresponding to each of the target distances are the same; determine the minimum sum of distances among the sums of the distances corresponding to the same first sample numbers; and determine each first sample number corresponding to each of the minimum sums of distances of the same second sample numbers as each of the first sample numbers corresponding to the second sample numbers; a receiving module, configured to receive each of the first sample numbers corresponding to each second sample number fed back by the server; the determining module, configured to determine the ordinary sample and the central sample corresponding to the second sample number as samples in the same cluster.

12. A device, characterized in that, The device includes a memory, a processor, and a clustering program stored in the memory and executable on the processor. When the clustering program is executed by the processor, it implements the steps of the sample clustering method according to any one of claims 1-4, or executes the steps of the sample clustering method according to any one of claims 5-9.

13. A readable storage medium, characterized in that the readable storage medium stores a clustering program, and when the clustering program is executed by a processor, it implements the steps of the sample clustering method according to any one of claims 1-9.

Citation Information

Patent Citations

  • User data similarity evaluation method and device, terminal equipment and storage medium

    CN110378749A

  • Spectral clustering method, device and system, computer equipment and storage medium

    CN111310817A