A data clustering method and related device

By introducing the first auxiliary array in the data clustering process, the problem of low computing efficiency in the prior art is solved, and more efficient data clustering is achieved.

CN111506730BActive Publication Date: 2025-07-29TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010305219.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-04-17
Publication Date
2025-07-29
Estimated Expiration
2040-04-17

AI Technical Summary

Technical Problem

The prior art has low computational efficiency in the process of data clustering and is difficult to meet current needs.

Method used

The first auxiliary array is introduced to reduce the calculation complexity and improve data clustering efficiency by counting the attractiveness and attribute degree in the similarity calculation between data samples.

Benefits of technology

By using the first auxiliary array, the complexity of calculating the degree of attribution is reduced and the efficiency of data clustering is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111506730B_ABST
    Figure CN111506730B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a data clustering method and related devices. During the data clustering process of N data samples by a computer device, multiple calculations and iterations of attraction and belonging degrees are performed to determine the data samples serving as clustering centers. In the t-th iteration among them, the attraction between N data samples can be calculated according to the similarity between the N data samples. During the process of calculating the attraction, a first auxiliary array corresponding to each of the N data samples can be statistically obtained. The first auxiliary array corresponding to the i-th data sample can reflect the sum of the effective attractions of the other data samples among the N data samples to the i-th data sample in the t-th iteration. This first auxiliary array can be applied to the calculation of the belonging degrees between N data samples in the t-th iteration, such that when calculating the belonging degree of a data sample relative to the i-th data sample, the first auxiliary array can be directly used. This method improves the efficiency of data clustering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing, and in particular, to a data clustering method and related devices. Background Art

[0002] Data clustering technology can classify data in various dimensions. For example, for text data samples, the text data samples can be divided into multiple categories according to text similarity, and the text similarity between text data samples in the same category is relatively high.

[0003] When performing data clustering through related technologies, it is necessary to calculate the attraction degree and belonging degree between samples through multiple iterations before determining the clustering center in the data samples, so as to achieve the final clustering.

[0004] The clustering calculation efficiency of the above-mentioned related technologies is relatively low, and it is difficult to meet the current data clustering requirements. Summary of the Invention

[0005] In order to solve the above technical problems, this application provides a data clustering method and related devices, which reduce the calculation complexity of calculating the belonging degree and improve the efficiency of data clustering.

[0006] The embodiments of this application disclose the following technical solutions:

[0007] On the one hand, the embodiments of this application provide a data clustering method. During the data clustering process of N data samples, it includes multiple calculation iterations for the attraction degree and belonging degree of the N data samples. The method is executed by a computer device, and the method includes:

[0008] In the t-th iteration, according to the similarity between N data samples, calculate the attraction degree between N data samples; the t-th iteration is one of the multiple calculation iterations;

[0009] Statistically calculate the first auxiliary array corresponding to each of the N data samples; among them, for the i-th data sample among the N data samples, the first auxiliary array corresponding to the i-th data sample is used to identify the sum of the effective attraction degrees of other data samples among the N data samples by the i-th data sample in the t-th iteration;

[0010] According to the attraction degree between N data samples and the first auxiliary array, calculate the belonging degree between N data samples in the t-th iteration.

[0011] On the other hand, the embodiments of this application provide a computer device, and the device includes a first calculation unit, a statistical unit, and a second calculation unit:

[0012] The first calculation unit is used for performing multiple calculation iterations on the attraction degree and belonging degree of N data samples during the data clustering process of the N data samples. In the t-th iteration, according to the similarity between the N data samples, the attraction degree between the N data samples is calculated; the t-th iteration is one of the multiple calculation iterations.

[0013] The statistical unit is used for statistically calculating the first auxiliary arrays respectively corresponding to the N data samples; wherein, for the i-th data sample among the N data samples, the first auxiliary array corresponding to the i-th data sample is used to identify the sum of the effective attraction degrees of the other data samples among the N data samples by the i-th data sample in the t-th iteration.

[0014] The second calculation unit is used for calculating the belonging degree between the N data samples in the t-th iteration according to the attraction degree between the N data samples and the first auxiliary arrays.

[0015] On the other hand, an embodiment of the present application provides a computer device, which includes a processor and a memory:

[0016] The memory is used for storing program codes and transmitting the program codes to the processor;

[0017] The processor is used for executing the above data clustering method according to the instructions in the program codes.

[0018] On the other hand, an embodiment of the present application provides a computer-readable storage medium, which is used for storing program codes, and the program codes are used for executing the above data clustering method.

[0019] It can be seen from the above technical solutions that during the data clustering process of N data samples by a computer device, multiple calculation iterations of the attraction degree and belonging degree are performed to determine the data samples serving as the clustering centers. In the t-th iteration among them, the attraction degree between the N data samples can be calculated according to the similarity between the N data samples. During the process of calculating the attraction degree, the first auxiliary arrays respectively corresponding to the N data samples can be statistically obtained. The first auxiliary array corresponding to the i-th data sample can reflect the sum of the effective attraction degrees of the other data samples among the N data samples by the i-th data sample in the t-th iteration. This first auxiliary array can be applied to the calculation of the belonging degree between the N data samples in the t-th iteration, so that when calculating the belonging degree of a data sample relative to the i-th data sample, the first auxiliary array can be directly used, without the need to additionally traverse the effective attraction degrees of other data samples relative to the i-th data sample, reducing the calculation complexity of calculating the belonging degree and improving the efficiency of data clustering. Description of the Drawings

[0020] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required in the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0021] Figure 1 Schematic diagram of an application scenario of a data clustering method provided by an embodiment of the present application;

[0022] Figure 2 Flowchart of a data clustering method provided by an embodiment of the present application;

[0023] Figure 3 Flowchart of an iterative method;

[0024] Figure 4 Schematic diagram of a concurrent computing mode scenario provided by an embodiment of the present application;

[0025] Figure 5 Schematic diagram of a concurrent computing mode scenario provided by an embodiment of the present application;

[0026] Figure 6a Structural diagram of a computer device provided by an embodiment of the present application;

[0027] Figure 6b Structural diagram of a computer device provided by an embodiment of the present application;

[0028] Figure 6c Structural diagram of a computer device provided by an embodiment of the present application;

[0029] Figure 7 Structural diagram of a computer device provided by an embodiment of the present application;

[0030] Figure 8 Structural diagram of a server provided by an embodiment of the present application. Detailed implementation manners

[0031] The following describes the embodiments of the present application in conjunction with the accompanying drawings.

[0032] When performing data clustering through related technologies, it is necessary to calculate the attraction and belonging degrees between data samples through multiple iterations before determining the clustering centers in the data samples and achieving the final clustering. In one iteration process of the related technology, when calculating the belonging degree of any data sample a relative to data sample b based on the attraction between data samples, in addition to referring to the attraction of data sample a relative to data sample b, it is also necessary to count the attractions of other data samples relative to data sample b. This results in the need to traverse a large amount of additional data when calculating the attraction between data samples, increasing the computational complexity, thus affecting the efficiency of the clustering calculation and making it difficult to meet the current data clustering requirements.

[0033] To this end, the embodiments of the present application provide a data clustering method, which introduces a first auxiliary array, so that when calculating the belonging degree of a data sample relative to a target data sample, there is no need to additionally traverse the attractions of other data samples relative to the target data sample, reducing the computational complexity of calculating the belonging degree and improving the efficiency of data clustering.

[0034] First, the execution subject of the embodiments of the present application is introduced. The data clustering method provided by the present application can be applied to a computer device, which can be a terminal device or a server. The terminal device can be, for example, a smart phone, a computer, a personal digital assistant (PDA), a tablet computer, a point of sales (POS), an in-vehicle computer, etc. The data clustering method can also be applied to a server, and the server can be an independent server, a server in a cluster, or a cloud server, etc.

[0035] In addition, it should be noted that the data clustering method provided by the embodiments of the present application can be applied to various scenarios that require clustering of data samples. In some embodiments, for example, it can be applied to the scenario of corpus clustering. In this scenario, the data samples to be clustered can be corpus samples, and the corpus samples are clustered through dimensions such as semantic similarity to obtain a set of samples with similar semantics in the corpus samples. Thus, functions such as generalizing and training the corpus based on the corpus samples with similar semantics and improving the performance of the semantic matching model can be realized.

[0036] For another example, the data clustering method can also be applied to the scenario of node clustering. In this scenario, the data samples to be clustered can be node samples, and the node samples can be data samples of different entities such as user data samples. By clustering the node samples, a set of nodes with similar entity characteristics can be obtained. For example, by clustering user data samples, a set of users with similar interests can be obtained, thus realizing functions such as friend recommendation.

[0037] Similarly, it can also be applied to various other scenarios that require clustering to achieve functions, such as image classification, question-and-answer systems for human-computer interaction, etc. By means of image clustering, semantic similarity mining, etc., the accuracy of image classification, the recall rate and accuracy of the question-and-answer system can be improved.

[0038] Next, taking the server as the execution subject and combining with the application scenario of semantic clustering, the data clustering method provided by the embodiments of the present application will be introduced.

[0039] See Figure 1 , Figure 1 which is a schematic diagram of the application scenario of a data clustering method provided by the embodiments of the present application. Figure 1 It includes a server 101. Assuming that there are N corpus samples to be clustered stored therein, based on the semantic information carried by the corpus samples, the server 101 can perform data clustering on the corpus samples by executing the data clustering method to obtain a sample set with close semantics.

[0040] The data clustering method provided by the embodiments of the present application can be an Affinity Propagation (AP) clustering method. This AP clustering method can be a method that calculates the clustering centers of data samples based on message passing between data samples, and then realizes data clustering.

[0041] Among them, during the data clustering process of the N corpus samples by the server 101, the similarity between the N corpus samples can be determined for calculating the attractiveness during the iterative process, and the similarity between the corpus samples does not change during the clustering process. In each iteration performed by the server 101, such as the t-th iteration, the attractiveness and the belonging degree between the N corpus samples can be calculated and updated. The attractiveness and the belonging degree are messages passed between corpus samples during the iterative process of the AP clustering method. Thus, after multiple iterations are completed, the clustering centers between the N corpus samples can be determined according to the attractiveness and the belonging degree between the corpus samples, and the clustering centers can be used to determine the corpus samples with close semantics to them, so as to realize semantic clustering according to the clustering centers.

[0042] To facilitate understanding of the technical solution provided by the embodiments of the present application, the relevant content of the similarity S, the attractiveness R, and the belonging degree A mentioned in this solution will be introduced below. The i-th corpus sample and the j-th corpus sample mentioned below belong to different corpus samples among the aforementioned N corpus samples, i = 1, 2, ……, N; j = 1, 2, ……, N.

[0043] S(i, j) is the similarity of the i-th corpus sample relative to the j-th corpus sample, which can reflect the ability of the j-th corpus sample to be the clustering center of the i-th corpus sample.

[0044] R(i,j) represents the attraction degree of the i-th corpus sample relative to the j-th corpus sample, which can reflect the degree to which the i-th corpus sample is passively attracted by the j-th corpus sample, that is, the objectivity degree of the j-th data sample becoming the clustering center of the i-th data sample. R(i,j) can be determined according to S(i,j).

[0045] A(i,j) represents the membership degree of the i-th corpus sample relative to the j-th corpus sample, which can reflect the tendency degree of the i-th corpus sample to actively select the j-th corpus sample as the clustering center, that is, the subjective degree of the j-th data sample becoming the clustering center of the i-th data sample.

[0046] After introducing the related contents such as similarity, attraction degree, and membership degree as above, the technical solution of this application will be introduced below.

[0047] In the embodiment of this application, referring to Figure 1 , in the t-th iteration, the server 101 can calculate R(i,j) according to the similarity S(i,j) between N corpus samples. It should be noted that unless otherwise specified, R(i,j) and A(i,j) mentioned below can refer to the data calculated in the t-th iteration.

[0048] In the related technology, when calculating A(i,j) between N corpus samples, as mentioned above, in addition to the R(i,j) of the i-th corpus sample relative to the j-th corpus sample, it is also necessary to traverse the attraction degrees of other corpus samples in these N corpus samples relative to the j-th corpus sample. Thus, in one iteration of the server 101, when calculating the total N 2 A(i,j) between N corpus samples, in addition to reading the R(i,j) between N corpus samples and other corpus samples, it is also necessary to perform an additional N 2 times of traversing the attraction degrees between these N corpus samples, resulting in the calculation complexity of the membership degree in one iteration being O(N 2 ).

[0049] To reduce the calculation complexity of the membership degree, in the embodiment of this application, in the t-th iteration, the server 101 can, during the process of calculating R(i,j), count the first auxiliary array Rsum(k) corresponding to each of the N corpus samples, where k = 1, 2,..., N. The Rsum(k) can be the data counted in the t-th round of iteration. The Rsum(k) can reflect the sum of the effective attraction degrees of other corpus samples in the N corpus samples relative to the k-th corpus sample in the t-th iteration, that is, Rsum(k) = ∑ i,i≠k max{0, R(i,k)}, and the effective attraction degree can be a non-negative attraction degree.

[0050] Thus, when the server 101 needs to calculate A(i,j), it only needs to read R(i,j) of the i-th corpus sample relative to the j-th corpus sample, and directly call Rsum(j), which is the sum of the effective attraction degrees of other corpus samples that have been statistically processed by the j-th corpus sample, without the need for additional traversal. Furthermore, in one iteration of the server 101, it only needs to read the attraction degrees R(i,j) between N corpus samples and other corpus samples and the first auxiliary array to calculate A(i,j) between the N corpus samples, and the computational complexity of this membership degree is O(N). It can be seen that this method reduces the computational complexity by an exponential level for the membership degree calculation method in each iteration, thereby improving the efficiency of data clustering.

[0051] Next, taking the server as the execution entity, the data clustering method provided by the embodiments of the present application will be introduced.

[0052] In this data clustering method, the number of data samples for clustering is N, where N is a positive integer. During the data clustering process of the N data samples, it includes multiple calculation iterations for the attraction degrees and membership degrees between the N data samples.

[0053] Before iteration, relevant parameters involved in the data clustering process can be preconfigured, including: reference degree p, damping coefficient λ, maximum number of iterations (maxiter), and convergence iteration number (Convergence_iter). Among them, p is the reference degree for a data sample to be a clustering center, used to adjust the clustering scale, and usually p is taken as the average value of the similarities between data samples; λ is used to alleviate the data oscillation phenomenon during iteration; Convergence_iter can be the threshold for the number of times the clustering center reaches a steady state calculated during the iteration process.

[0054] See Figure 2 , which shows a flowchart of a data clustering method provided by the embodiments of the present application. As Figure 2 shown, the method includes:

[0055] S201: In the t-th iteration, calculate the attraction degrees between N data samples according to the similarities between the N data samples.

[0056] Where the t-th iteration is any one of the above multiple calculation iterations.

[0057] In the embodiments of the present application, based on the fact that there is similarity information between some of the N data samples in certain aspects, thus, the similarities S(i,j) between the N data samples can be calculated.

[0058] The embodiments of the present application do not limit the expression form of the similarity S(i,j). According to the actual scenario or different requirements, a suitable way can be selected to express the similarity S(i,j). For example, the similarity between two data samples can be expressed in the form of the distance between the two data samples. The similarity determined in this way is in a symmetric form, that is, S(i,j) = S(j,i). The similarity between data samples can also be expressed in other forms, and the obtained similarity can also be in an asymmetric form, that is, S(i,j) ≠ S(j,i).

[0059] It should be noted that in the embodiments of the present application, S(i,i), R(i,i), and A(i,i) between data samples themselves can be calculated. Among them, S(i,i) can be the initial degree of the i-th data sample as a clustering center, and S(i,i) can be set as the reference degree p. R(i,i) reflects the objective evaluation level of the i-th data sample as a clustering center for itself, and A(i,i) can reflect the confidence degree of the i-th data sample as a clustering center for itself. In addition, before the first iteration, R and A between N data samples can be initialized. For example, R and A between data samples are initialized to 0.

[0060] It should be noted that during multiple iterations, the similarity between these data samples remains fixed and will not change due to the iteration process.

[0061] In the t-th iteration, R between N data samples can be calculated according to S between these N data samples.

[0062] S202: Statistically calculate the first auxiliary arrays corresponding to N data samples respectively.

[0063] Among them, the first auxiliary array Rsum(i) corresponding to the i-th data sample among N data samples can be used to identify the sum of the effective attraction degrees of other data samples among N data samples relative to the i-th data sample in the t-th iteration, Rsum(i) = ∑ j,j≠i max{0, R(j,i)}. The effective attraction degree mentioned here can be a non-negative attraction degree, and the effective attraction degree of the i-th data sample relative to the j-th data sample is max{0, R(i,j)}.

[0064] In this way, in the embodiments of the present application, the first auxiliary arrays corresponding to N data samples can be statistically calculated during the process of calculating the attraction degrees between N data samples.

[0065] S203: Calculate the membership degrees between N data samples in the t-th iteration according to the attraction degrees between N data samples and the first auxiliary arrays.

[0066] In a specific implementation, the A(i, j) between N data samples in the t-th iteration can be calculated according to the following formulas (1) and (2):

[0067]

[0068] A(i, j) = λA′(i, j) + (1 - λ)A t-1 (i, j); (2)

[0069] The A t-1 (i, j) in the above formula (2) can be the membership degree of the i-th data sample relative to the j-th data sample in the (t - 1)-th iteration.

[0070] Through the above method, the membership degrees between N data samples in the t-th iteration are calculated.

[0071] It can be seen from the above technical solution that during the data clustering process of N data samples by a computer device, multiple calculations and iterations of the attraction degree and membership degree are performed to determine the data samples serving as the clustering centers. In the t-th iteration among them, the attraction degree between N data samples can be calculated according to the similarity between N data samples. During the calculation of the attraction degree, a first auxiliary array corresponding to each of the N data samples can be statistically obtained. The first auxiliary array corresponding to the i-th data sample can be reflected in the t-th iteration as the sum of the effective attraction degrees of other data samples in the N data samples to the i-th data sample. This first auxiliary array can be applied to the calculation of the membership degrees between N data samples in the t-th iteration, so that when calculating the membership degree of a data sample relative to the i-th data sample, the first auxiliary array can be directly used, without the need to additionally traverse the effective attraction degrees of other data samples relative to the i-th data sample, reducing the computational complexity of calculating the membership degree and improving the efficiency of data clustering.

[0072] In the related art, when calculating R(i, j) in the t-th iteration, in addition to referring to the S(i, j) of the i-th data sample relative to the j-th data sample, it is also necessary to traverse the similarities between the i-th data sample and other data samples, as well as the membership degrees between the i-th data sample and other data samples in the previous (t - 1)-th iteration, to obtain the maximum value of the sum of the similarities and membership degrees between the i-th data sample and other data samples. That is to say, when calculating the total N 2 R(i, j) between N data samples in the t-th iteration, it is necessary to additionally traverse the similarities between data samples N 2 times and the membership degrees between data samples in the previous iteration. The computational complexity of calculating R(i, j) between data samples in the related art is O(N 2 ), which affects the efficiency of data clustering.

[0073] Therefore, in a possible implementation manner, during the process of calculating the membership degrees of N data samples in the t-th iteration in S203, the method further includes: determining second auxiliary arrays respectively corresponding to the N data samples according to the membership degrees.

[0074] Among them, the second auxiliary array corresponding to the i-th data sample among the N data samples can be used to identify the maximum membership information corresponding to the i-th data sample and the data sample corresponding to the maximum membership information in the t-th iteration.

[0075] The maximum membership information corresponding to the i-th data sample can be denoted as ASmax1(i), and ASmax1(i) can refer to the maximum value among the sum of the similarity degrees and membership degrees of the i-th data sample relative to other data samples (including the i-th data sample). The data sample corresponding to the maximum membership information can be denoted as the ASmax1 index (i)-th data sample. Among them, the sum of the similarity degree and membership degree of the i-th data sample relative to the ASmax1 index (i)-th data sample is the maximum value among the sum of the similarity degrees and membership degrees of the i-th data sample relative to other data samples. Among them, ASmax1(i) and ASmax1 index (i) can be data statistically obtained in the t-th iteration and used for calculating the attraction degree in the (t + 1)-th iteration.

[0076] That is to say, during the process of calculating A(i, j) between data samples in the t-th iteration, the maximum membership information ASmax1(i) corresponding to the i-th data sample and the data sample corresponding to the maximum membership information, that is, the ASmax1 index (i)-th data sample, can be statistically obtained.

[0077] In this way, when calculating R(i, j) between N data samples in the (t + 1)-th iteration, if the j-th data sample does not belong to the ASmax1 index (i)-th data sample statistically obtained in the t-th iteration (that is, j ≠ ASmax1 index (i)), the similarity degree between the i-th data sample and the j-th data sample can be read, and the statistically obtained ASmax1(i) can be directly called, and R(i, j) in the (t + 1)-th iteration can be calculated through the following formulas (3) and (4). There is no need to additionally traverse the membership degrees of the i-th data sample relative to other data samples in the t-th iteration.

[0078]

[0079] R(i, j) = λR’(i, j) + (1 - λ)R t-1 (i, j); (4)

[0080] The R in the above formula (4) t-1 (i, j) can be the attraction degree between the i-th data sample and the j-th data sample in the (t - 1)-th iteration.

[0081] Furthermore, in one iteration, only the S(i, j) of N corpus samples relative to other corpus samples and the second auxiliary array need to be read to complete the calculation of R(i, j), and the computational complexity of calculating this attraction degree is reduced to O(N). This method reduces an exponential computational complexity of calculating the attraction degree and improves the efficiency of data clustering.

[0082] In addition, in a possible implementation manner, during the process of calculating the membership degrees of N data samples in the t-th iteration in S203, the method further includes: determining the third auxiliary array corresponding to each of the N data samples according to the membership degrees.

[0083] Among them, the third auxiliary array corresponding to the i-th data sample is used to identify the second-largest membership information corresponding to the i-th data sample in the t-th iteration. The second-largest membership information corresponding to the i-th data sample may refer to the second-largest value (i.e., the second-largest value) among the sum of the similarity degrees and membership degrees between the i-th data sample and other data samples (including the i-th data sample), denoted as ASmax2(i).

[0084] That is to say, during the process of calculating A(i, j) between data samples in the t-th iteration, ASmax2(i) can be counted.

[0085] In this way, when calculating R(i, j) between N data samples in the (t + 1)-th iteration, if the j-th data sample is the ASmax1 index (i)-th data sample statistically obtained in the t-th iteration (i.e., j = ASmax1 index (i)), the similarity degree between the i-th data sample and the j-th data sample can be read, and the statistically obtained ASmax2(i) can be called, and through the above formulas (3) and (4), the attraction degree R(i, j) in the (t + 1)-th iteration can be calculated. There is no need to additionally traverse the membership degrees of the i-th data sample relative to other data samples.

[0086] In this method, when calculating R(i, j) between data samples in the (t + 1)-th iteration, the computational complexity of calculating the attraction degree of the i-th data sample relative to the ASmax1 index (i)-th data sample is reduced, thereby further improving the efficiency of data clustering.

[0087] It should be noted that the embodiments of the present application do not limit the N 2The similarity, as well as the representation methods of the attraction degree and the belonging degree in each iteration. According to the actual situation or different requirements, a suitable method can be selected to represent these data among the data samples.

[0088] In a possible implementation, the similarity S, the attraction degree R, and the belonging degree A among the data samples can be represented in the form of a matrix.

[0089] The similarity among the N data samples can be represented by a similarity matrix, such as the similarity matrix (1) shown below, where the i-th row is used to reflect the S(i,j) between the i-th data sample and other data samples (including the i-th data sample). For example, in this similarity matrix, the first row reflects the similarities between the first data sample and other data samples (including the first data sample), which are S(1,1), S(1,j), …, S(1,N) respectively.

[0090]

[0091] Similarly, as shown in the attraction matrix (2) below, the attraction degree among the N data samples can be represented by the attraction matrix, where the i-th row can be used to reflect the R(i,j) between the i-th data sample and other data samples (including the i-th data sample).

[0092]

[0093] As shown in the belonging degree matrix (3) below, the belonging degree of the N data samples can be represented by the belonging degree matrix, where the i-th row is used to reflect the A(i,j) between the i-th data sample and other data samples (including the i-th data sample).

[0094]

[0095] Representing the similarity, the attraction degree, and the belonging degree among the N data samples in the form of a matrix facilitates data traversal and improves the efficiency of data reading during the iteration process.

[0096] It can be understood that in an actual scenario, some data samples may be relatively similar in some aspects, while some data samples may be completely dissimilar. Based on the fact that the messages in the AP clustering algorithm in this application are globally transmitted, the similarity between some data samples may be relatively high, and the similarity between some data samples may be very low. When S(i,j) of the i-th data sample relative to the j-th data sample is close to 0, R(i,j) calculated by the above formulas (3)-(4) is <0, and the corresponding effective attractiveness is 0. In this case, whether or not R(i,j) is calculated, it will not affect the value of the first auxiliary array Rsum(j) of the j-th data sample. Additionally, when S(i,j) is close to 0, A(i,j) calculated by the above formulas (1)-(2) is ≤0, so that A(i,j)+S(i,j) is close to 0. Thus, whether or not A(i,j) is calculated, it will not affect the maximum membership information ASmax1(i) in the second auxiliary array corresponding to the i-th data sample, nor will it affect the second-largest membership information ASmax2(i) corresponding to the third auxiliary array.

[0097] That is to say, when S(i,j) is low, whether or not to calculate the corresponding R(i,j) and A(i,j) will not affect the results of the attractiveness and membership degrees between other data samples. And for the case where S(i,j) is low, since the corresponding R(i,j) and A(i,j) are both low, the corresponding data samples will not become clustering centers.

[0098] It can be seen that ignoring the calculation of R and A for data samples with extremely low similarity will not affect the final data clustering result. The message passing calculation process for such data samples is redundant.

[0099] Therefore, in a possible implementation manner, the method further includes:

[0100] S301: Determine the target positions in the similarity matrix where the similarity is less than a threshold.

[0101] Among them, the threshold can be used to determine the similarity for which the calculation of attractiveness and membership degree is not required. When it is determined that the similarity is less than this threshold, the calculation of the attractiveness and membership degree of the corresponding data samples can be omitted.

[0102] In the embodiments of this application, when the similarity in the similarity matrix is less than the threshold, the position of this similarity in the similarity matrix can be determined as the target position.

[0103] S302: Set the values at the target positions in the similarity matrix, the attractiveness matrix, and the membership degree matrix to be empty.

[0104] It can be understood that the data at the target positions in the similarity matrix, attraction matrix, and membership matrix all correspond to the same data sample. That is, assuming the target position is the \(i\)-th row and \(j\)-th column, the data at this target position in these three matrices are respectively the similarity, attraction, and membership of the \(i\)-th data sample relative to the \(j\)-th data sample.

[0105] Among them, the attraction matrix and membership matrix can be the attraction matrix and membership matrix in each iteration.

[0106] After setting the values at the target positions in the similarity matrix, attraction matrix, and membership matrix to empty, it is achieved to avoid the iterative calculation of \(R\) and \(A\) for this part of the data at the target positions.

[0107] In a specific implementation, after converting the original dense matrix (a matrix with a relatively large proportion of non-zero elements) into a sparse matrix (a matrix with a relatively large proportion of zero elements), the matrix can be stored in the Compressed Storage Row (CSR) format.

[0108] In this method, by setting the target positions in the matrix to empty, the dense matrix is converted into a sparse matrix, so that the space complexity is reduced from \(O(N 2 ) to \(O(M)\), where \(M\) is the number of non-zero elements in the matrix after data processing, saving the memory space for matrix storage. If \(M\ll N 2 , the storage space can be greatly reduced. In addition, in each iteration, the calculation of the attraction and membership at the target positions in the matrix is avoided, reducing the computational time complexity to \(O(M)\) and improving the efficiency of data clustering.

[0109] In the related technology, for the case where \(S\), \(R\), and \(A\) between data samples are represented in matrix form, see Figure 3 , which shows a flowchart of an iterative method. As Figure 3 shown, when calculating \(R(i, j)\) of the \(i\)-th data sample relative to other data samples in the \(t\)-th iteration, in addition to referring to \(S(i, j)\) of the \(i\)-th data sample relative to other data samples, that is, the row data of the \(i\)-th data sample in the similarity matrix (framed by the gray line box), it is also necessary to traverse the membership of the \(i\)-th data sample relative to other data samples in the previous (\(t - 1\))-th iteration, that is, the row data of the \(i\)-th data sample in the membership matrix of the (\(t - 1\))-th iteration (framed by the gray line box), to obtain the maximum value of the sum of the similarity and membership of the \(i\)-th data sample relative to other data samples, and then achieve the attraction calculation.

[0110] That is to say, when calculating the attraction degree of the i-th data sample relative to other data samples in the t-th iteration, it can be obtained by traversing the row data corresponding to the i-th data sample in the similarity matrix and the row data of the i-th data sample in the membership degree matrix in the (t - 1)-th iteration.

[0111] When calculating A(i, j) of the i-th data sample relative to the j-th data sample, it is necessary to traverse the column data of the j-th column in the attraction degree matrix of this iteration, that is, the R(i, j) of other data samples relative to the j-th data sample (framed by the gray line box), to obtain the sum of the effective attraction degrees of other data samples relative to the j-th data sample, so as to realize the calculation of the membership degree. For example, when calculating A(i, 1), it is necessary to traverse the column data of the 1st column in this attraction degree matrix; when calculating A(i, N), it is necessary to traverse the column data of the Nth column in this attraction degree matrix. That is to say, when calculating the membership degree of the i-th data sample relative to other data samples, it is necessary to traverse not only the row data corresponding to the i-th data sample in the attraction degree matrix, but also each column data in this attraction degree matrix.

[0112] Based on the fact that when calculating A(i, j) of the i-th data sample relative to other data samples in each iteration, it is necessary to traverse each column data of the complete attraction degree matrix, which results in the inability to split the attraction degree matrix to perform concurrent calculation of the membership degree of some data samples relative to other data samples.

[0113] In an embodiment of the present application, in a possible implementation manner, a method for concurrently calculating the attraction degree and the membership degree in each iteration is provided.

[0114] First of all, it should be noted that the three matrices used for concurrent calculation can be the dense matrices that have not undergone data processing as described above, or the sparse matrices that have undergone data processing as described above. The present application does not make any limitations in this regard.

[0115] In an embodiment of the present application, the N data samples can be divided into M sample groups, where M < N. Each sample group can include one or more data samples, so as to realize concurrent calculation of these M sample groups respectively. Among them, the concurrent calculation mentioned in the present application can refer to a method of splitting the data into multiple parts and performing parallel calculation of each part in different threads. By concurrently calculating A and R corresponding to the data samples in each data group, the efficiency of calculating A and R between all data samples can be improved.

[0116] Thus, in the t-th iteration, the method for calculating the attraction degrees between N data samples according to the similarities between N data samples in S201 above includes:

[0117] Calculate the attraction degrees of M sample groups concurrently. Among them, the attraction degrees between the data samples in the sample group and N data samples can be calculated based on the row data corresponding to the data samples in the similarity matrix.

[0118] It can be understood that in the t-th iteration, when calculating the attraction degree of a data sample (denoted as the i-th data sample) in the sample group relative to other data samples, the similarity between this data sample and other data samples can be obtained through the row data corresponding to the i-th data sample in the similarity matrix, so as to calculate the attraction degree of the i-th data sample relative to other data samples.

[0119] In S203, calculate the attribution degrees between N data samples in the t-th iteration according to the attraction degrees between N data samples and the first auxiliary array, including:

[0120] In the t-th iteration, calculate the attribution degrees of M sample groups concurrently. Among them, the attribution degrees between the data samples in the sample group and N data samples can be calculated based on the row data corresponding to the data samples in the attraction matrix and the first auxiliary array.

[0121] Since the first auxiliary array corresponding to the data samples in each sample group has been pre-statistically calculated, and this first auxiliary array is the sum of the effective attraction degrees of other data samples relative to this data sample. Therefore, when calculating the attribution degree of the i-th data sample in the sample group relative to the j-th data sample, directly call the first auxiliary array corresponding to the j-th data sample for attribution degree calculation, without having to determine the sum of the effective attraction degrees of other data samples for the j-th data sample through column traversal.

[0122] That is to say, by introducing the first auxiliary array, it is no longer necessary to perform column traversal for each column of the attraction matrix when calculating the attribution degree of a data sample relative to other data samples.

[0123] In this method, when calculating the attraction degree and attribution degree of the i-th data sample relative to other data samples, only by using the first auxiliary array corresponding to these N data samples, and performing row traversal on the row data corresponding to the i-th data sample in the similarity matrix, attraction matrix and attribution matrix, the calculation of attraction degree and attribution degree can be completed, without the need for other data in the matrix, thus making it possible to perform concurrent calculation of the attraction degree and attribution degree between data samples.

[0124] It should be noted that the embodiments of the present application do not limit the number of concurrent calculations for the attractiveness and attribution. In specific implementations, it can be determined according to the computing power of the computer device. On a multi-core processor, several thread pools can be created, and the similarity matrix, attractiveness matrix, and attribution matrix can be evenly distributed to each thread to perform concurrent calculations on the attractiveness, attribution, and the first auxiliary array, the second auxiliary array, and the third auxiliary array.

[0125] The following is an example of this concurrent calculation method. Refer to Figure 4 , which shows a schematic diagram of a scenario of a concurrent calculation method provided by an embodiment of the present application. As Figure 4 shown, in the t-th iteration, assuming there are 100 data samples, the corresponding similarity matrix and the attribution matrix in the (t - 1)-th iteration are 100 * 100 matrices. Among them, the 100 data samples can be divided into 3 (i.e., M = 3) sample groups, namely sample group 1, sample group 2, and sample group 3. Sample group 1 includes the first 30 data samples; sample group 2 includes the 31st to 70th data samples; sample group 3 includes the last 30 data samples.

[0126] Then, as Figure 4 shown in the attractiveness calculation process corresponding to the thin solid line, the second auxiliary array (ASmax1(1)-ASmax1(100) and ASmax1 index (1)-ASmax1 index (100)) and the third auxiliary array (ASmax2(1)-ASmax2(100)) corresponding to each data sample can be determined according to the similarity matrix and the attribution matrix in the (t - 1)-th iteration. The row data corresponding to the first 30 data samples of the similarity matrix, and the corresponding second auxiliary array and third auxiliary array (ASmax1(1)-ASmax1(30), ASmax1 index (1)-ASmax1 index (30) and ASmax2(1)-ASmax2(30)) can be input into thread 1 to calculate the attractiveness of the first 30 data samples relative to the 100 data samples.

[0127] Similarly, the row data corresponding to the 31st to 70th data samples of the similarity matrix, and the corresponding second auxiliary array and third auxiliary array (ASmax1(31)-ASmax1(70), ASmax1 index (31)-ASmax1 index(70) and ASmax2(31) - ASmax2(70)) are input into thread 2 to calculate the attraction degrees of the 31st to 70th data samples relative to the 100 data samples. The row data corresponding to the last 30 data samples of the similarity matrix and the corresponding second auxiliary array and third auxiliary array (ASmax1(71) - ASmax1(100), ASmax1 index (71) - ASmax1 index (100) and ASmax2(71) - ASmax2(100)) are input into thread 3 to calculate the attraction degrees of the last 30 data samples relative to the 100 data samples.

[0128] After completing the calculation of the attraction degrees among the 100 data samples, the first auxiliary array (Rsum(1) - Rsum(100)) corresponding to each data sample can be determined accordingly.

[0129] As Figure 4 shown in the membership degree calculation process corresponding to the thick solid line, the attraction degrees of the first 30 data samples relative to other data samples and the corresponding first auxiliary array (Rsum(1) - Rsum(30)) are input into thread 1 to calculate the membership degrees of the first 30 data samples relative to the 100 data samples. Similarly, the attraction degrees of the 31st to 70th data samples relative to other data samples and the corresponding first auxiliary array (Rsum(31) - Rsum(70)) are input into thread 2 to calculate the membership degrees of the 31st to 70th data samples relative to the 100 data samples. The attraction degrees of the last 30 data samples and the corresponding first auxiliary array (Rsum(71) - Rsum(100)) are input into thread 3 to calculate the membership degrees of the last 30 data samples relative to the 100 data samples.

[0130] In this method, by introducing the first auxiliary array, the column traversal for membership degree calculation against matrix data is avoided, which provides the possibility for concurrent calculation, thereby reducing the iterative calculation time through concurrent calculation and improving the data clustering efficiency.

[0131] In an embodiment of the present application, after completing multiple calculation iterations for N data samples, in a possible implementation manner, the method further includes:

[0132] S401: Determine the target data sample serving as the clustering center according to the attraction degrees and membership degrees among the N data samples.

[0133] In an embodiment of the present application, when the number of iterations reaches the maximum number of iterations, or when the consecutive number of clustering centers determined in each iteration reaches the convergence iteration number, the target data sample serving as the clustering center is determined according to the attraction degrees and membership degrees among the N data samples.

[0134] In an embodiment of the present application, for example, the actually existing data samples corresponding to R(i, i)+A(i, i)>0 can be determined as clustering centers, and the data samples serving as the clustering centers can be determined as target data samples.

[0135] S402: Complete data clustering for N data samples according to the target data samples.

[0136] In a specific implementation, the method of completing data clustering according to target data samples includes: if a data sample is not connected to any target data sample, the data sample can be classified independently. The connection between the data samples mentioned here can mean that the similarity between the two is relatively high. If a data sample is connected to a unique target data sample, it is determined that the data sample belongs to the class where the target data sample is located. If a data sample is connected to multiple target data samples, it is determined that the data sample belongs to the class where the target data sample with the greatest similarity among all the connected target data samples is located.

[0137] It can be understood that when this solution is applied to the scenario of corpus clustering, according to the methods of S401-S402 above, target data samples can be determined from corpus samples after multiple calculation iterations, and clustering for all corpus samples can be completed according to the target data samples, so as to obtain a set of samples with similar semantics in the corpus samples. Thus, the training corpus can be generalized according to the corpus samples with similar semantics to improve functions such as the performance of the semantic matching model.

[0138] When applied to the scenario of node clustering, the target data samples serving as clustering centers can be determined from node samples through the methods of S401-S402, and clustering for all node samples can be achieved based on this, and a set of node samples with similar entity characteristics can be determined, so as to implement functions such as friend recommendation according to these node sample sets.

[0139] Similarly, when applied to other various scenarios that need to implement functions through clustering, clustering for all data samples can be achieved through the above methods. For example, image classification, question-and-answer systems for human-computer interaction, etc., so as to improve the accuracy of image classification, the recall rate and accuracy of the question-and-answer system through image clustering, semantic similarity mining, etc.

[0140] In this way, clustering for N data samples is achieved.

[0141] Next, the data clustering method provided in the embodiment of the present application will be introduced in combination with actual application scenarios.

[0142] This embodiment provides a sparse affinity propagation clustering method that supports multi-threaded concurrent computing and can quickly cluster data samples according to the similarity degree between data samples. In this method, by abstracting data samples into network nodes, after inputting the similarity sparse matrix between data samples, based on the similarity between network nodes, two types of messages, namely the attractiveness and the responsibility between network nodes, can be calculated concurrently in multi-threads in each iteration. Through multiple rounds of iteration, the attractiveness and the responsibility of each network node are continuously updated until several high-quality clustering centers, that is, target data samples, are generated, so as to allocate all data samples to the corresponding clusters.

[0143] See Figure 5 , which shows a flowchart of a data clustering method provided by an embodiment of the present application. As Figure 5 shown, the method includes:

[0144] S501: Construct a corresponding similarity sparse matrix according to N data samples to be clustered.

[0145] Since there is similarity information in certain aspects between data samples, a similarity matrix between the N data samples can be constructed, and data processing is performed on the similarity matrix in the foregoing manner to obtain a corresponding similarity sparse matrix.

[0146] S502: Configure relevant parameters and initialize the responsibility matrix, the attractiveness matrix, and the auxiliary arrays.

[0147] Among them, the auxiliary arrays may include a first auxiliary array, a second auxiliary array, and a third auxiliary array.

[0148] Before the first iteration, the responsibility matrix, the attractiveness matrix, and the auxiliary arrays can be initialized, for example, making all the data in them be 0.

[0149] Regarding the parameters that need to be configured as described above, they will not be elaborated here.

[0150] S503: Concurrently update the responsibility matrix, the attractiveness matrix, and the auxiliary arrays, and calculate the target data samples serving as clustering centers.

[0151] In the embodiment of the present application, in each iteration, the responsibility, the attractiveness, and the auxiliary arrays between data samples can be updated in a concurrent computing manner to calculate the target data samples serving as clustering centers in each iteration.

[0152] S504: Determine whether the clustering iteration end condition is satisfied.

[0153] If so, execute S505; if not, execute S503.

[0154] Among them, the end condition of the clustering iteration is that the number of iterations reaches the maximum number of iterations, or the consecutive number of clustering centers determined in each iteration reaches the convergence iteration number.

[0155] S505: Determine the class to which the data sample belongs.

[0156] Among them, the specific way to determine the class to which the data sample belongs is as described in the previous S402, and will not be elaborated here.

[0157] In the embodiments of the present application, based on 8.6k corpus samples and the semantic similarity based on 29.8M similarity relationships, and on the premise of the same configuration parameters, the experimental results shown in Table 1 are obtained. Through the comparison of the experimental results, the solution of the multi-threaded concurrent computing mode of the present application is superior to other related technologies in terms of memory overhead and clustering efficiency.

[0158] Table 1 Experimental Results

[0159]

[0160]

[0161] Based on the data clustering method provided in the foregoing embodiments, the embodiments of the present application provide a computer device, which may be the computer device mentioned above. Refer to Figure 6a , which shows a structural diagram of a computer device provided in the embodiments of the present application. The device includes a first calculation unit 601, a statistical unit 602, and a second calculation unit 603:

[0162] The first calculation unit 601 is used to perform multiple calculation iterations on the attraction and belonging degrees of the N data samples during the data clustering process of the N data samples. In the t-th iteration, according to the similarity between the N data samples, calculate the attraction between the N data samples; the t-th iteration is one of the multiple calculation iterations;

[0163] The statistical unit 602 is used to count the first auxiliary arrays corresponding to the N data samples respectively; among them, for the i-th data sample among the N data samples, the first auxiliary array corresponding to the i-th data sample is used to identify the sum of the effective attraction degrees of the other data samples among the N data samples by the i-th data sample in the t-th iteration;

[0164] The second calculation unit 603 is used to calculate the belonging degrees between the N data samples in the t-th iteration according to the attraction between the N data samples and the first auxiliary array.

[0165] In a possible implementation manner, the second calculation unit 603 is specifically used for:

[0166] In the process of calculating the membership degrees of N data samples in the t-th iteration, second auxiliary arrays corresponding to the N data samples are determined according to the membership degrees; wherein, the second auxiliary array corresponding to the i-th data sample is used to identify the maximum membership information corresponding to the i-th data sample and the data sample corresponding to the maximum membership information in the t-th iteration; the second auxiliary array is used to calculate the attraction degrees between the N data samples in the (t + 1)-th iteration.

[0167] In a possible implementation manner, the second calculation unit 603 is specifically configured to:

[0168] In the process of calculating the membership degrees of N data samples in the t-th iteration, third auxiliary arrays corresponding to the N data samples are determined according to the membership degrees; wherein, the third auxiliary array corresponding to the i-th data sample is used to identify the second largest membership information corresponding to the i-th data sample in the t-th iteration; the third auxiliary array is used to calculate the attraction degrees between the N data samples in the (t + 1)-th iteration.

[0169] In a possible implementation manner, the similarity degrees between the N data samples are represented by a similarity matrix, where the i-th row is used to reflect the similarity degrees between the i-th data sample and other data samples;

[0170] The attraction degrees between the N data samples are represented by an attraction matrix, where the i-th row is used to reflect the attraction degrees between the i-th data sample and other data samples;

[0171] The membership degrees of the N data samples are represented by a membership matrix, where the i-th row is used to reflect the membership degrees between the i-th data sample and other data samples.

[0172] In a possible implementation manner, the first calculation unit 601 is specifically configured to:

[0173] The N data samples are divided into M sample groups, M < N. In the t-th iteration, the attraction degrees are calculated for the M sample groups concurrently; wherein, according to the row data corresponding to the data samples in the sample groups in the similarity matrix, the attraction degrees between the data samples in the sample groups and the N data samples are calculated;

[0174] The second calculation unit 603 is specifically configured to:

[0175] In the t-th iteration, the membership degrees are calculated for the M sample groups concurrently; wherein, according to the row data corresponding to the data samples in the sample groups in the attraction matrix and the first auxiliary array, the membership degrees between the data samples in the sample groups and the N data samples are calculated.

[0176] In a possible implementation manner, refer to Figure 6b, the figure shows a structural diagram of a computer device provided by an embodiment of the present application. The device further includes a setting unit 604, and the setting unit 604 is specifically configured to:

[0177] Determine the target positions in the similarity matrix where the similarity is less than the threshold;

[0178] Set the values at the target positions in the similarity matrix, the attraction matrix, and the membership matrix to be empty.

[0179] In a possible implementation, refer to Figure 6c , the figure shows a structural diagram of a computer device provided by an embodiment of the present application. The device further includes a clustering unit 605, and the clustering unit 605 is specifically configured to:

[0180] Determine the target data samples serving as the clustering centers according to the attraction and membership among N data samples;

[0181] Complete data clustering for the N data samples according to the target data samples.

[0182] It can be seen from the above technical solutions that during the data clustering process of the computer device for N data samples, multiple calculation iterations of attraction and membership are performed to determine the data samples serving as the clustering centers. In the t-th iteration, the attraction among the N data samples can be calculated according to the similarity among the N data samples. During the calculation of the attraction, a first auxiliary array corresponding to each of the N data samples can be statistically obtained. The first auxiliary array corresponding to the i-th data sample can reflect the sum of the effective attractions of other data samples among the N data samples to the i-th data sample in the t-th iteration. This first auxiliary array can be applied to the calculation of the membership among the N data samples in the t-th iteration, so that when calculating the membership of a data sample relative to the i-th data sample, the first auxiliary array can be directly used without having to additionally traverse the effective attractions of other data samples relative to the i-th data sample, reducing the computational complexity of calculating the membership and improving the efficiency of data clustering.

[0183] An embodiment of the present application also provides a computer device. The computer device will be introduced below with reference to the accompanying drawings. Please refer to Figure 7 As shown, an embodiment of the present application provides a structural diagram of a computer device. The device 700 may also be a terminal device. Taking the terminal device as a mobile phone as an example:

[0184] Figure 7 What is shown is a partial structural block diagram of the mobile phone provided by an embodiment of the present application. Refer to Figure 7, the mobile phone includes components such as a Radio Frequency (RF) circuit 710, a memory 720, an input unit 730, a display unit 740, sensors 750, an audio circuit 760, a wireless fidelity (WiFi) module 770, a processor 780, and a power supply 790. Those skilled in the art can understand that Figure 7 the mobile phone structure shown in

[0185] does not limit the mobile phone and may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. Figure 7 The following specifically introduces each component of the mobile phone:

[0186] The RF circuit 710 can be used for receiving and sending signals during information reception or call processes. Specifically, after receiving the downlink information from the base station, it is given to the processor 780 for processing; in addition, it sends the designed uplink data to the base station. Generally, the RF circuit 710 includes but is not limited to antennas, at least one amplifier, a transceiver, a coupler, a Low Noise Amplifier (LNA), a duplexer, etc. In addition, the RF circuit 710 can also communicate with the network and other devices through wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.

[0187] The memory 720 can be used to store software programs and modules. The processor 780 executes various functional applications and data processing of the mobile phone by running the software programs and modules stored in the memory 720. The memory 720 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 720 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0188] The input unit 730 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function control of the mobile phone. Specifically, the input unit 730 may include a touch panel 731 and other input devices 732. The touch panel 731, also known as a touch screen, can collect touch operations of the user on or near it (such as operations of the user using a finger, a stylus, or any suitable object or accessory on or near the touch panel 731), and drive corresponding connection devices according to a preset program. Optionally, the touch panel 731 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 780, and can receive commands sent by the processor 780 and execute them. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch panel 731. In addition to the touch panel 731, the input unit 730 may further include other input devices 732. Specifically, the other input devices 732 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, etc.

[0189] The display unit 740 can be used to display information input by the user or information provided to the user, as well as various menus of the mobile phone. The display unit 740 may include a display panel 741. Optionally, the display panel 741 can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. Further, the touch panel 731 can cover the display panel 741. When the touch panel 731 detects a touch operation on or near it, it is transmitted to the processor 780 to determine the type of touch event. Subsequently, the processor 780 provides a corresponding visual output on the display panel 741 according to the type of touch event. Although in Figure 7 the touch panel 731 and the display panel 741 are implemented as two independent components to realize the input and input functions of the mobile phone, in some embodiments, the touch panel 731 and the display panel 741 can be integrated to realize the input and output functions of the mobile phone.

[0190] The mobile phone may further include at least one sensor 750, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 741 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 741 and / or the backlight when the mobile phone is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity, and can be used for applications that identify the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors such as gyroscope, barometer, hygrometer, thermometer, infrared sensor that the mobile phone can also be configured with, they will not be elaborated here.

[0191] The audio circuit 760, the speaker 761, and the microphone 762 can provide an audio interface between the user and the mobile phone. The audio circuit 760 can transmit the electrical signal converted from the received audio data to the speaker 761, and the speaker 761 converts it into a sound signal for output; on the other hand, the microphone 762 converts the collected sound signal into an electrical signal, which is received by the audio circuit 760 and converted into audio data. After the audio data is output to the processor 780 for processing, it is sent to another mobile phone through the RF circuit 710, for example, or the audio data is output to the memory 720 for further processing.

[0192] WiFi belongs to short-distance wireless transmission technology. The mobile phone can help users send and receive emails, browse the web, and access streaming media through the WiFi module 770. It provides users with wireless broadband Internet access. Although Figure 7The WiFi module 770 is shown, but it can be understood that it does not belong to the essential components of the mobile phone and can be omitted entirely within the scope of not changing the essence of the invention as needed.

[0193] The processor 780 is the control center of the mobile phone, connecting various parts of the entire mobile phone using various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 720, and by calling data stored in the memory 720, it performs various functions of the mobile phone and processes data. Optionally, the processor 780 may include one or more processing units; preferably, the processor 780 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 780 either.

[0194] The mobile phone also includes a power supply 790 (such as a battery) for powering each component. Preferably, the power supply can be logically connected to the processor 780 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system.

[0195] Although not shown, the mobile phone may also include a camera, a Bluetooth module, etc., which will not be elaborated here.

[0196] In this embodiment, the processor 780 included in the terminal device further has the following functions:

[0197] During the data clustering process of N data samples, it includes multiple calculation iterations for the attractiveness and belongingness of the N data samples.

[0198] In the t-th iteration, according to the similarity between the N data samples, calculate the attractiveness between the N data samples; the t-th iteration is one of the multiple calculation iterations.

[0199] Statistically calculate the first auxiliary array corresponding to each of the N data samples; among them, for the i-th data sample among the N data samples, the first auxiliary array corresponding to the i-th data sample is used to identify the sum of the effective attractiveness of other data samples among the N data samples by the i-th data sample in the t-th iteration.

[0200] According to the attractiveness between the N data samples and the first auxiliary array, calculate the belongingness between the N data samples in the t-th iteration.

[0201] The computer device provided in the embodiment of the present application may be a server, please refer to Figure 8 as shown Figure 8This is a structural diagram of the server 800 provided by the embodiments of the present application. The server 800 may vary greatly due to different configurations or performances, and may include one or more central processing units (CPUs) 822 (for example, one or more processors) and a memory 832, and one or more storage media 830 (for example, one or more mass storage devices) for storing application programs 842 or data 844. Among them, the memory 832 and the storage media 830 may be transient storage or persistent storage. The program stored in the storage media 830 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Further, the central processing unit 822 may be configured to communicate with the storage media 830 and execute a series of instruction operations in the storage media 830 on the server 800.

[0202] The server 800 may further include one or more power supplies 826, one or more wired or wireless network interfaces 850, one or more input / output interfaces 858, and / or one or more operating systems 841, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and so on.

[0203] The steps in the above embodiments may also be executed by a server, and the server may be based on the Figure 8 server structure shown.

[0204] The embodiments of the present application further provide a computer-readable storage medium, which is used to store program codes, and the program codes are used to execute the methods described in the foregoing embodiments.

[0205] The embodiments of the present application further provide a computer program product including instructions, which, when running on a computer, cause the computer to execute the methods described in the foregoing embodiments.

[0206] In the description of this application and the above-mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described herein can be implemented in an order different from those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0207] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist simultaneously. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0208] In several embodiments provided by this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0209] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0210] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0211] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0212] As described above, the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application.

[0213] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium can be at least one of the following media: read-only memory (ROM), RAM, magnetic disks, or optical discs and other media that can store program codes.

[0214] It should be noted that the various embodiments in this specification are described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the device and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the description of the method embodiments. The device and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.

[0215] As described above, this is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A data clustering method, characterized in that, During the data clustering process of N data samples, it includes multiple calculation iterations for the attraction and belonging degrees of the N data samples. The method is executed by a computer device, and the data samples include corpus samples. The method includes: In the t-th iteration, according to the similarity between N data samples, calculate the attraction between N data samples; the t-th iteration is one of the multiple calculation iterations; Statistically calculate the first auxiliary array corresponding to each of the N data samples; among them, for the i-th data sample among the N data samples, the first auxiliary array corresponding to the i-th data sample is used to identify the sum of the effective attraction of other data samples among the N data samples by the i-th data sample in the t-th iteration; According to the attraction between N data samples and the first auxiliary array, calculate the belonging degree between N data samples in the t-th iteration.

2. The method according to claim 1, characterized in that, During the process of calculating the belonging degree between N data samples in the t-th iteration, the method further includes: Determine the second auxiliary array corresponding to each of the N data samples according to the belonging degree; among them, the second auxiliary array corresponding to the i-th data sample is used to identify the maximum belonging information corresponding to the i-th data sample and the data sample corresponding to the maximum belonging information in the t-th iteration; the second auxiliary array is used to calculate the attraction between N data samples in the (t + 1)-th iteration.

3. The method according to claim 1, wherein During the process of calculating the belonging degree between N data samples in the t-th iteration, the method further includes: Determine the third auxiliary array corresponding to each of the N data samples according to the belonging degree; among them, the third auxiliary array corresponding to the i-th data sample is used to identify the second-largest belonging information corresponding to the i-th data sample in the t-th iteration; the third auxiliary array is used to calculate the attraction between N data samples in the (t + 1)-th iteration.

4. The method according to any one of claims 1 to 3, characterized in that, The similarity between the N data samples is represented by a similarity matrix, where the i-th row is used to reflect the similarity between the i-th data sample and other data samples; The attraction between N data samples is represented by an attraction matrix, where the i-th row is used to reflect the attraction between the i-th data sample and other data samples; The belonging degree of N data samples is represented by a belonging degree matrix, where the i-th row is used to reflect the belonging degree between the i-th data sample and other data samples.

5. The method according to claim 4, wherein The N data samples are divided into M sample groups, M < N. The calculating the attraction between N data samples according to the similarity between N data samples includes: In the t-th iteration, calculate the attraction for the M sample groups concurrently; among them, according to the row data corresponding to the data samples in the sample group in the similarity matrix, calculate the attraction between the data samples in the sample group and the N data samples; The calculating the belonging degree between N data samples in the t-th iteration according to the attraction between N data samples and the first auxiliary array includes: In the t-th iteration, calculate the belonging degree for the M sample groups concurrently; among them, according to the row data corresponding to the data samples in the sample group in the attraction matrix and the first auxiliary array, calculate the belonging degree between the data samples in the sample group and the N data samples.

6. The method according to claim 4, wherein The method further includes: Determining target positions in the similarity matrix where the similarity is less than a threshold; Setting the values at the target positions in the similarity matrix, the attraction matrix, and the membership matrix to be empty.

7. The method according to claim 1, characterized in that After completing multiple calculation iterations, the method further includes: Determining target data samples as clustering centers according to the attraction and membership among N data samples; Completing data clustering for the N data samples based on the target data samples.

8. A computer device, characterized in that, The device includes a first calculation unit, a statistical unit, and a second calculation unit: The first calculation unit is configured to, during the data clustering process of N data samples, including multiple calculation iterations for the attraction and membership of the N data samples, calculate the attraction among the N data samples according to the similarity among the N data samples in the t-th iteration; the t-th iteration is one of the multiple calculation iterations, and the data samples include corpus samples; The statistical unit is configured to count the first auxiliary arrays respectively corresponding to the N data samples; wherein, for the i-th data sample among the N data samples, the first auxiliary array corresponding to the i-th data sample is used to identify the sum of the effective attraction of the other data samples among the N data samples by the i-th data sample in the t-th iteration; The second calculation unit is configured to calculate the membership among the N data samples in the t-th iteration according to the attraction among the N data samples and the first auxiliary array.

9. The device according to claim 8, characterized in that, The second calculation unit is specifically configured to: During the process of calculating the membership of the N data samples in the t-th iteration, determine the second auxiliary arrays respectively corresponding to the N data samples according to the membership; wherein, the second auxiliary array corresponding to the i-th data sample is used to identify the maximum membership information corresponding to the i-th data sample and the data sample corresponding to the maximum membership information in the t-th iteration; the second auxiliary array is used to calculate the attraction among the N data samples in the (t + 1)-th iteration.

10. The device according to claim 8, characterized in that, The second calculation unit is specifically configured to: During the process of calculating the membership of the N data samples in the t-th iteration, determine the third auxiliary arrays respectively corresponding to the N data samples according to the membership; wherein, the third auxiliary array corresponding to the i-th data sample is used to identify the second-largest membership information corresponding to the i-th data sample in the t-th iteration; the third auxiliary array is used to calculate the attraction among the N data samples in the (t + 1)-th iteration.

11. The device according to any one of claims 8 to 10, characterized in that The similarity among the N data samples is represented by a similarity matrix, where the i-th row is used to reflect the similarity between the i-th data sample and other data samples; The attraction among the N data samples is represented by an attraction matrix, where the i-th row is used to reflect the attraction between the i-th data sample and other data samples; The membership of the N data samples is represented by a membership matrix, where the i-th row is used to reflect the membership between the i-th data sample and other data samples.

12. The device according to claim 11, characterized in that, The first calculation unit is specifically configured to: N data samples are divided into M sample groups, where M < N. In the t-th iteration, the attractiveness of the M sample groups is calculated concurrently. Among them, according to the row data corresponding to the data samples in the sample group in the similarity matrix, the attractiveness between the data samples in the sample group and the N data samples is calculated. The second calculation unit is specifically used for: In the t-th iteration, the belongingness of the M sample groups is calculated concurrently. Among them, according to the row data corresponding to the data samples in the sample group in the attractiveness matrix and the first auxiliary array, the belongingness between the data samples in the sample group and the N data samples is calculated.

13. The device according to claim 11, characterized in that, The device further includes a setting unit, and the setting unit is specifically used for: Determine the target positions in the similarity matrix where the similarity is less than the threshold. Set the values at the target positions in the similarity matrix, the attractiveness matrix, and the belongingness matrix to be empty.

14. The device according to claim 8, wherein, The device further includes a clustering unit, and the clustering unit is specifically used for: Determine the target data samples serving as the clustering centers according to the attractiveness and belongingness among the N data samples. Complete the data clustering for the N data samples according to the target data samples.

15. A computer device, characterized in that, The device includes a processor and a memory: The memory is used to store program codes and transmit the program codes to the processor. The processor is used to execute the data clustering method according to any one of claims 1-7 based on the instructions in the program codes.

16. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store program codes, and the program codes are used to execute the data clustering method according to any one of claims 1-7.

17. A computer program product, characterized in that, The computer program product includes instructions, and when the instructions run on a computer device, the computer device is caused to execute the data clustering method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Model training method, device and equipment

    CN109800798A

  • Affinity propagation clustering method based on genetic algorithm

    CN110543913A