Sample point clustering method, device, processor and electronic device
Patent Information
- Application Number
- CN202310652938.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-02
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2043-06-02
AI Technical Summary
[0004]本申请的主要目的在于提供一种样本点的聚类方法、装置和处理器及电子设备,以解决相关技术中聚类的准确性较低的问题
[0019]This application employs the following steps: K sample points are obtained from a sample set containing N sample points, and these K sample points are determined as K initial centroids, where N is a positive integer greater than 2 and K is a positive integer less than N; based on these K initial centroids, other sample points in the sample set are clustered to obtain K first sample subsets centered on each initial centroid, wherein the sample points in each first sample subset are sorted from closest to farthest according to their cosine distance from the corresponding initial centroid; edge selection is performed on each of the K first sample subsets to obtain K second sample subsets, wherein each second sample subset includes multiple sample points from the corresponding first sample subset, with preset weights and the furthest sorted points; based on all the sample points included in each second sample subset and other K-1 second sample subsets... For each centroid associated with the subset, the (K+1)th centroid is determined, and the (K+1)th centroid is associated with the (K+1)th first sample subset. Under the condition of satisfying a preset convergence condition, based on the determined centroid, the other sample points in the sample set (excluding the determined centroid) are clustered to obtain the clustered target sample set. Based on the initial centroid and the corresponding initial clustering results, and based on edge selection and weighting of the initial clustering results, new centroids are obtained. Based on edge selection and weighting, the new centroids are uniformly distributed with the initial centroid. Through continuous iteration, a final number of uniformly distributed centroids are obtained, solving the problem of overly close or scattered distribution caused by random centroid selection, ensuring the uniformity of centroid selection, and thus achieving the technical effect of improving the accuracy of clustering.
Smart Images

Figure CN116701981B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of big data, and more specifically, to a method, apparatus, processor, and electronic device for clustering sample points. Background Technology
[0002] Clustering algorithms need to determine cluster centers, i.e., centroids (or centroid points), during the initialization phase. Existing technologies often use random sampling to determine a preset number of centroids. However, this method may result in centroids that are too scattered or too concentrated, leading to uneven centroid distribution and thus lower accuracy in subsequent clustering.
[0003] There is currently no effective solution to the problem of low clustering accuracy in related technologies. Summary of the Invention
[0004] The main objective of this application is to provide a clustering method, apparatus, processor, and electronic device for sample points to solve the problem of low clustering accuracy in related technologies.
[0005] To achieve the above objectives, according to one aspect of this application, a clustering method for sample points is provided. The method includes: obtaining K sample points from a sample set comprising N sample points, and determining the K sample points as K initial centroids, where N is a positive integer greater than 2 and K is a positive integer less than N; clustering other sample points in the sample set based on the K initial centroids to obtain K first sample subsets centered at each initial centroid, wherein the sample points in each first sample subset are sorted from nearest to farthest according to their cosine distance to the corresponding initial centroid; and performing edge selection processing on the K first sample subsets respectively to obtain K... The second sample subset includes multiple sample points with preset weights and the furthest order from the corresponding first sample subset. Based on all the sample points included in each second sample subset and the centroids associated with the other K-1 second sample subsets, the K+1th centroid is determined, wherein the K+1th centroid is associated with K+1 first sample subsets. Under the condition of satisfying the preset convergence condition, based on the determined centroids, the other sample points in the above sample set except for the determined centroids are clustered to obtain the clustered target sample set.
[0006] As an optional approach, determining the K+1th centroid based on all sample points included in each second sample subset and the centroids associated with each of the other K-1 second sample subsets includes: determining representative sample points associated with each second sample subset based on the average weighted distance of all sample points included in each second sample subset and the K-1 initial centroids associated with the other K-1 second sample subsets, wherein the representative sample points are the sample points with the largest average weighted distance among all second sample subsets; determining the target sample point with the largest average weighted distance from all K of the above representative sample points, and determining the target sample point as the additional centroid, thus obtaining the above K+1 centroids.
[0007] As an optional approach, determining the representative sample points associated with each second sample subset includes: determining the current i-th second sample subset from the K second sample subsets, and obtaining the average weighted distance between each sample point in the i-th second sample subset and the K-1 initial centroids; wherein, obtaining the average weighted distance between the m-th sample point in the i-th second sample subset and the K-1 initial centroids includes: obtaining the cosine distances between the m-th sample point and each centroid of the K-1 initial centroids, obtaining K-1 first cosine distance values; obtaining the K-1... The product of each cosine distance value and the number of sample points in the second sample subset containing the associated centroid is used to obtain K-1 second cosine distance values. The third cosine distance value obtained by summing the above K-1 second cosine distance values is determined by the quotient of K-1 and the third cosine distance value, which is the average weighted distance between the above m-th sample point and the above K-1 initial centroids. The j-th sample point included in the above i-th second sample subset is determined as the representative sample point associated with the above i-th sample subset, wherein the average weighted distance of the above j-th sample point is the largest among all sample points included in the above i-th second sample subset.
[0008] As an optional approach, after determining the (K+1)th centroid, the method further includes: clustering the other sample points in the sample set based on the (K+1)th centroid to obtain (K+1)th third sample subsets, wherein the sample points in each third sample subset are sorted from nearest to farthest according to the cosine distance between them and the corresponding centroid; performing edge selection processing on the (K+1)th third sample subsets to obtain (K+1)th fourth sample subsets, wherein each fourth sample subset includes multiple sample points with preset weights and the furthest sorted points within the corresponding third sample subset; determining the (K+2)th centroid based on all the sample points included in each fourth sample subset and the centroids associated with the other K fourth sample subsets, wherein the obtained K+2 centroids are associated with K+2 third sample subsets; repeating the above steps until the preset convergence condition is met.
[0009] As an optional approach, satisfying the preset convergence condition includes: determining that the preset convergence condition is satisfied when the number of centroids determined above reaches a preset number threshold; or determining that the preset convergence condition is satisfied when the cosine distance between each sample point in each of the obtained second sample subsets and the corresponding centroid is not greater than a preset cosine distance.
[0010] As an optional approach, obtaining K sample points from a sample set containing N sample points and determining the K sample points as K initial centroids includes: traversing to obtain the cosine distance between each pair of the N sample points; selecting at least two sample points with the largest cosine distance, and determining the at least two sample points as the K sample points.
[0011] As an optional approach, before obtaining K sample points from the sample set containing N sample points, the method further includes: obtaining the sample set from a historical error log set, wherein the sample set includes error log data of different error categories, and the historical error log set includes error log data generated by the client-associated account during historical usage periods; after obtaining the clustered target sample set, the method further includes: performing feature extraction processing on the target sample set to obtain the erroneous operation behavior features of the account, wherein the erroneous operation behavior features are used to indicate the erroneous operation habits of the account during the historical time period.
[0012] To achieve the above objectives, according to another aspect of this application, a clustering apparatus for sample points is provided. The apparatus includes: an acquisition unit, configured to acquire K sample points from a sample set comprising N sample points, and determine the K sample points as K initial centroids, wherein N is a positive integer greater than 2, and K is a positive integer less than N; a first clustering unit, configured to cluster other sample points in the sample set based on the K initial centroids, obtaining K first sample subsets centered at each initial centroid, wherein the sample points in each first sample subset are sorted from closest to farthest according to the cosine distance between them and their corresponding initial centroids; and an edge selection unit, configured to perform edge selection processing on each of the K first sample subsets. The system obtains K second sample subsets, where each second sample subset includes multiple sample points from the corresponding first sample subset with preset weights and the furthest order. A determination unit is used to determine the (K+1)th centroid based on all sample points included in each second sample subset and the centroids associated with the other K-1 second sample subsets, where the obtained K+1 centroids are associated with K+1 first sample subsets. A second clustering unit is used, under the condition of satisfying preset convergence, to cluster the other sample points in the above sample set besides the determined centroids, to obtain the clustered target sample set.
[0013] As an optional scheme, the above-mentioned determining unit includes: a first determining module, used to determine representative sample points associated with each second sample subset based on the average weighted distance of all sample points included in each second sample subset to K-1 initial centroid points associated with other K-1 second sample subsets, wherein the representative sample point is the sample point with the largest average weighted distance among all second sample subsets; and a second determining module, used to determine the target sample point with the largest average weighted distance from all K representative sample points, and determine the target sample point as an additional centroid point, thereby obtaining the above-mentioned K+1 centroid points.
[0014] As an optional solution, the first determining module includes: a first determining submodule, used to determine the current i-th second sample subset from the K second sample subsets, and obtain the average weighted distance between each sample point in the i-th second sample subset and the K-1 initial centroids; wherein, obtaining the average weighted distance between the m-th sample point in the i-th second sample subset and the K-1 initial centroids includes: a first obtaining submodule, used to obtain the cosine distance between the m-th sample point and each centroid of the K-1 initial centroids, to obtain K-1 first cosine distance values; a second obtaining submodule, used to obtain the K-1 first cosine distance values. The product of each cosine distance in the cosine distance values and the number of sample points in the second sample subset where the associated centroid point is located is used to obtain K-1 second cosine distance values; the second determining submodule is used to determine the quotient of the third cosine distance value obtained by summing the above K-1 second cosine distance values and K-1 as the average weighted distance between the above m-th sample point and the above K-1 initial centroid points; the third determining submodule is used to determine the representative sample point associated with the above i-th sample subset by the j-th sample point included in the above i-th sample subset, wherein the average weighted distance of the above j-th sample point is the largest among all sample points included in the above i-th second sample subset.
[0015] As an optional solution, the above-mentioned apparatus further includes: a clustering module, used to cluster other sample points in the sample set based on the K+1 centroid points after determining the K+1 centroid points, to obtain K+1 third sample subsets, wherein the sample points in each third sample subset are sorted from near to far according to the cosine distance between them and the corresponding centroid points; and an edge selection module, used to perform edge selection processing on the K+1 third sample subsets after determining the K+1 centroid points, to obtain K+1 fourth sample subsets. In this process, each fourth sample subset includes multiple sample points from the corresponding third sample subset, with preset weights and the furthest order. The third determination module is used to determine the K+2th centroid point after determining the K+1th centroid point, based on all sample points included in each fourth sample subset and the centroid points associated with the other K fourth sample subsets. The obtained K+2 centroid points are associated with K+2 third sample subsets. The execution module is used to repeat the above steps after determining the K+1th centroid point until the preset convergence condition is met.
[0016] As an optional solution, the above-mentioned device further includes: a fourth determining module, used to determine that the preset convergence condition is met when the number of the determined centroids reaches a preset number threshold; or a fifth determining module, used to determine that the preset convergence condition is met when the cosine distance between each sample point in each of the obtained second sample subsets and the corresponding centroid is not greater than a preset cosine distance.
[0017] As an optional solution, the above acquisition unit includes: a first acquisition module, used to traverse and acquire the cosine distance between each pair of the above N sample points; and a sixth determination module, used to select at least two sample points with the largest cosine distance, and determine the above at least two sample points as the above K sample points.
[0018] As an optional solution, the above-mentioned apparatus further includes: a second acquisition module acquiring the sample set from a historical error log set before acquiring K sample points from the sample set including N sample points, wherein the sample set includes error log data of different error categories, and the historical error log set includes error log data generated by the client-associated account during its use in a historical time period; the above-mentioned apparatus further includes: a feature extraction module, used to perform feature extraction processing on the target sample set after obtaining the clustered target sample set to obtain the erroneous operation behavior features of the account, wherein the erroneous operation behavior features are used to indicate the erroneous operation habits of the account in the historical time period.
[0019] This application employs the following steps: K sample points are obtained from a sample set containing N sample points, and these K sample points are determined as K initial centroids, where N is a positive integer greater than 2 and K is a positive integer less than N; based on these K initial centroids, other sample points in the sample set are clustered to obtain K first sample subsets centered on each initial centroid, wherein the sample points in each first sample subset are sorted from closest to farthest according to their cosine distance from the corresponding initial centroid; edge selection is performed on each of the K first sample subsets to obtain K second sample subsets, wherein each second sample subset includes multiple sample points from the corresponding first sample subset, with preset weights and the furthest sorted points; based on all the sample points included in each second sample subset and other K-1 second sample subsets... For each centroid associated with the subset, the (K+1)th centroid is determined, and the (K+1)th centroid is associated with the (K+1)th first sample subset. Under the condition of satisfying a preset convergence condition, based on the determined centroid, the other sample points in the sample set (excluding the determined centroid) are clustered to obtain the clustered target sample set. Based on the initial centroid and the corresponding initial clustering results, and based on edge selection and weighting of the initial clustering results, new centroids are obtained. Based on edge selection and weighting, the new centroids are uniformly distributed with the initial centroid. Through continuous iteration, a final number of uniformly distributed centroids are obtained, solving the problem of overly close or scattered distribution caused by random centroid selection, ensuring the uniformity of centroid selection, and thus achieving the technical effect of improving the accuracy of clustering. Attached Figure Description
[0020] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:
[0021] Figure 1 This is a flowchart of a clustering method for sample points provided in the embodiments of this application;
[0022] Figure 2 This is a schematic diagram of a clustering method for sample points provided in an embodiment of this application;
[0023] Figure 3 This is a schematic diagram of a clustering method for sample points provided in an embodiment of this application;
[0024] Figure 4 This is a schematic diagram of a clustering method for sample points provided in an embodiment of this application;
[0025] Figure 5This is a schematic diagram of a clustering device for sample points provided according to an embodiment of this application;
[0026] Figure 6 This is a schematic diagram of a clustering electronic device for sample points provided in the embodiments of this application. Detailed Implementation
[0027] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0028] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this application described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0030] It should be noted that all information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) involved in this disclosure are information and data authorized by the user or fully authorized by all parties. For example, this system has an interface with relevant users or organizations. Before obtaining relevant information, it is necessary to send an acquisition request to the aforementioned user or organization through the interface, and obtain the relevant information after receiving consent information from the aforementioned user or organization.
[0031] The present invention will now be described in conjunction with preferred implementation steps. Figure 1 This is a flowchart of the Z method provided according to an embodiment of this application, such as... Figure 1 As shown, the method includes the following steps:
[0032] Step S101: Obtain K sample points from the sample set including N sample points, and determine the K sample points as K initial centroids, where N is a positive integer greater than 2 and K is a positive integer less than N;
[0033] Step S102: Based on the K initial centroids, cluster the other sample points in the sample set to obtain K first sample subsets centered on each initial centroid. The sample points in each first sample subset are sorted from near to far according to the cosine distance between them and the corresponding initial centroid.
[0034] Step S103: Perform edge selection processing on the K first sample subsets respectively to obtain K second sample subsets, wherein each second sample subset includes multiple sample points with preset weights and the furthest sorting from the corresponding first sample subset;
[0035] Step S104: Based on all the sample points included in each second sample subset and the centroids associated with each of the other K-1 second sample subsets, determine the K+1th centroid, wherein the obtained K+1 centroids are associated with K+1 first sample subsets.
[0036] Step S105: If the preset convergence condition is met, cluster the other sample points in the sample set except for the determined centroid points based on the determined centroid points to obtain the clustered target sample set.
[0037] Optionally, in this embodiment, the above-mentioned clustering method for sample points can be applied, but is not limited to, log clustering scenarios. Using this sample point clustering method, it is possible not only to find similar logs and analyze errors that occur when user-associated accounts use related applications, but also to group logs with the same error type into the same group, and to analyze and mine logs in the same group to identify potential errors in user habits, and provide corresponding suggestions.
[0038] Furthermore, compared to the existing technology of directly randomly selecting a preset number of centroids for clustering, this embodiment uses the above-mentioned clustering method based on the initial centroids and the corresponding initial clustering results. Based on the edge selection and weighting of the initial clustering results, new centroids are obtained. The new centroids obtained based on edge selection and weighting are "uniformly distributed" with the initial centroids. Through continuous iteration, a final number of uniformly distributed centroids are obtained, avoiding the problem of overly close or scattered distribution caused by random selection of centroids. This ensures the uniformity of centroid selection and thus improves the accuracy of subsequent centroid-based clustering algorithms.
[0039] Optionally, in this embodiment, the sample set may include, but is not limited to, N sample points, wherein each sample point may be used to indicate a log data, the log data may be, but is not limited to, the record information generated by the user's associated account during the use of the relevant application, and may be used to indicate usage error information, where N is a positive integer greater than 2.
[0040] Optionally, in this embodiment, the K initial centroids may be, but are not limited to, the K sample points obtained from the sample set that are far apart by cosine distance. In the case of K being 2, the K initial centroids may be, but are not limited to, the two sample points in the sample set that are far apart by cosine distance. K is a positive integer less than N.
[0041] Optionally, in this embodiment, after determining K initial centroids, the other (NK) sample points in the sample set are clustered with each initial centroid as the center to obtain K first sample subsets. Each of the other (NK) sample points may, but is not limited to, select the initial centroid that is closest to its own sample point cosine distance and belong to the corresponding first sample subset.
[0042] Optionally, in this embodiment, the sample points included in the first sample subset are sorted according to the cosine distance between them and the corresponding centroid points. The centroid points with closer cosine distances are arranged in the relatively inner region of the first sample subset, and the centroid points with farther cosine distances are arranged in the relatively outer region of the first sample subset.
[0043] Optionally, in this embodiment, the edge selection process may be used, but is not limited to, to indicate the acquisition of multiple sample points with preset weights and the furthest order from the edge portion (relative to the outer region) of the first sample subset.
[0044] To illustrate further, with a preset weight of 10%, 10% of the sample points in the last 90%-100% region of the first sample subset are selected to obtain the second sample subset, wherein the number of sample points included in the second sample subset is 10% of the number of sample points included in the first sample subset.
[0045] Optionally, in this embodiment, after determining the second sample subsets corresponding to the K first sample subsets respectively, the K+1th centroid is determined based on all sample points included in each second sample subset and each centroid point associated with the other K-1 second sample subsets, wherein the obtained K+1 centroid points are associated with K+1 first sample subsets.
[0046] It should be noted that after determining the (K+1)th centroid, and if the preset convergence condition is not met, the (K+1)th centroid can be used as (K+1) new initial centroids of the sample set, and clustering can be performed based on these (K+1) new initial centroids to obtain (K+1) third sample subsets centered on each new initial centroid. Edge selection processing can then be performed on these (K+1) third sample subsets to obtain corresponding (K+1) fourth sample subsets. Finally, based on all sample points included in each third sample subset and the centroids associated with the other (K-1) fourth sample subsets, the (K+2)th centroid is determined…
[0047] Optionally, in this embodiment, the preset convergence condition may include, but is not limited to, the determination of centroid points reaching a preset number threshold, and may also include, but is not limited to, the cosine distance between each sample point of each obtained second sample subset and the corresponding centroid point being less than or equal to the preset cosine distance.
[0048] It should be noted that, under the condition of satisfying the preset convergence condition, based on the determined centroid, the other sample points in the sample set other than the determined centroid are clustered to obtain the clustered target sample set.
[0049] It should be noted that after obtaining the target sample set after clustering, the target sample set can be used, but is not limited to, to analyze the erroneous habit characteristics of user accounts associated with the sample data over historical time periods.
[0050] To further illustrate, such as Figure 2 As shown, an optional clustering method for sample points includes:
[0051] Step S201: Determine two initial centroids with the largest cosine distance from the set of 8 sample points (corresponding to...). Figure 2 (a) The initial two black spheres), and based on the two initial centroids, other sample points (corresponding to) Figure 2 Cluster the other 6 white balls in (a) (the dashed line represents the set of centroids to which the sample points of the white ball belong after clustering) to obtain two initial cluster sets;
[0052] Step S202: Based on a preset weight (50%), edge selection is performed on the two initial cluster sets to obtain the outer half of the sample points of each initial cluster set, forming two edge cluster sets (corresponding to...). Figure 2 (The sample points included in the two outer ring regions of (c));
[0053] Step S203: Based on the sample points and other centroids included in each edge set, determine the third centroid from all sample points included in the two edge cluster sets (e.g., ...). Figure 2 The three centroids (d) are used to re-cluster the other five sample points based on the three centroids, resulting in three new initial cluster sets corresponding to the three centroids.
[0054] It should be noted that after obtaining the new initial cluster sets corresponding to the three centroids, it is possible, but not limited to, to further obtain the three new edge cluster sets corresponding to the three initial cluster sets, and to determine the fourth centroid based on the relationship between the sample points included in each new edge set and other centroids... and so on until the number of determined centroids reaches the requirement or the sample points in each clustered set reach the clustering effect.
[0055] The embodiments provided in this application involve obtaining K sample points from a sample set containing N sample points and determining these K sample points as K initial centroids, where N is a positive integer greater than 2 and K is a positive integer less than N; based on the K initial centroids, the other sample points in the sample set are clustered to obtain K first sample subsets centered on each initial centroid, wherein the sample points in each first sample subset are sorted from closest to farthest according to the cosine distance between them and the corresponding initial centroid; edge selection processing is then performed on each of the K first sample subsets to obtain K... Each second sample subset includes multiple sample points from the corresponding first sample subset, with preset weights and the furthest order. Based on all sample points included in each second sample subset and the centroids associated with the other K-1 second sample subsets, the K+1th centroid is determined, where the K+1th centroid is associated with K+1 first sample subsets. Under preset convergence conditions, based on the determined centroids, the other sample points in the sample set (excluding the determined centroids) are clustered to obtain the clustered target sample set. Based on the initial centroids and the corresponding initial clustering results, and based on edge selection and weighting of the initial clustering results, new centroids are obtained. Based on edge selection and weighting, the new centroids are evenly distributed with the initial centroids. Through continuous iteration, a final number of evenly distributed centroids are obtained, solving the problem of overly close or scattered distribution caused by random centroid selection, ensuring the uniformity of centroid selection, and thus improving the accuracy of clustering.
[0056] As an optional approach, based on all sample points included in each second sample subset and the centroids associated with the other K-1 second sample subsets, the K+1th centroid is determined to include:
[0057] S1, based on the average weighted distance of all sample points included in each second sample subset to the K-1 initial centroids associated with the other K-1 second sample subsets, determine the representative sample point associated with each second sample subset, where the representative sample point is the sample point with the largest average weighted distance among all second sample subsets;
[0058] S2, determine the target sample point with the largest average weighted distance from all K representative sample points, and determine the target sample point as the additional centroid point, thus obtaining K+1 centroid points.
[0059] Optionally, in this embodiment, the average weighted distance may be, but is not limited to, the average weighted value of the cosine distance between the current sample point of the current second sample subset and the K-1 initial centroid points associated with the other K-1 second sample subsets.
[0060] To further illustrate, taking the second sample subset C1 as an example, the average weighted distance of the m-th sample point C1_m in C1 is: dist_C1_m = (cosine distance between C1_m and the centroid of C2 * total number of sample points in C2 + cosine distance between C1_m and the centroid of C3 * total number of sample points in C3 + ... + cosine distance between C1_m and the centroid of Ck-1 * total number of sample points in Ck-1) / (K-1), where C2, C3, and Ck-1 are the other K-1 second sample subsets.
[0061] Optionally, in this embodiment, the average weighted distance of all sample points included in each second sample subset is obtained, and the representative sample point with the largest average weighted distance in each second sample subset is determined, wherein there are K representative sample points in total.
[0062] Optionally, in this embodiment, the sample point with the largest average weighted distance is further obtained from the K representative sample points to obtain the target sample point, wherein the target sample point is the (K+1)th centroid point.
[0063] Through the embodiments provided in this application, representative sample points associated with each second sample subset are determined based on the average weighted distance of all sample points included in each second sample subset to the K-1 initial centroids associated with other K-1 second sample subsets. The representative sample point is the sample point with the largest average weighted distance among all second sample subsets. From all K representative sample points, the target sample point with the largest average weighted distance is determined and designated as an additional centroid, resulting in K+1 centroids. Based on the initial centroids and the corresponding initial clustering results, and based on edge selection and weighting of the initial clustering results, new centroids are obtained. The new centroids obtained based on edge selection and weighting have a "uniform distribution relationship" with the initial centroids. Through continuous iteration, a final number of uniformly distributed centroids are obtained, solving the problem of overly close or scattered distribution caused by random selection of centroids, ensuring the uniformity of centroid selection, and thus achieving the technical effect of improving the accuracy of clustering.
[0064] As an optional approach, the representative sample points associated with each second sample subset are determined as follows:
[0065] S1, determine the current i-th second sample subset from the K second sample subsets, and obtain the average weighted distance between each sample point and the K-1 initial centroids in the i-th second sample subset;
[0066] The average weighted distance between the m-th sample point in the i-th second sample subset and the K-1 initial centroid points includes:
[0067] S2, obtain the cosine distances between the m-th sample point and each of the K-1 initial centroids, and obtain the K-1 first cosine distance values;
[0068] S3, obtain the product of each cosine distance in the K-1 cosine distance values and the number of sample points in the second sample subset where the associated centroid point is located, to obtain K-1 second cosine distance values;
[0069] S4, the quotient of the third cosine distance obtained by summing K-1 second cosine distance values and K-1 is determined as the average weighted distance between the m-th sample point and the K-1 initial centroid points;
[0070] S5, determine the representative sample point associated with the i-th sample subset by the j-th sample point included in the i-th second sample subset, wherein the average weighted distance of the j-th sample point is the largest among all sample points included in the i-th second sample subset.
[0071] Optionally, in this embodiment, all sample points included in each second sample subset are traversed, the average weighted distance of each sample point to the other K-1 initial centroids is obtained, and the sample point with the largest average weighted distance in each second sample subset is determined as the representative sample point, and the sample point with the largest average weighted distance from all representative sample points is determined as the new centroid.
[0072] Optionally, in this embodiment, the first cosine distance value may be used, but is not limited to, to indicate the cosine distance between the sample point and the corresponding initial centroid point; the second cosine distance value may be used, but is not limited to, to indicate the product of the above cosine distance and the number of sample points in the second sample subset where the corresponding initial centroid point is located; and the average weighted distance may be used, but is not limited to, to indicate the sum of the K-1 second cosine distance values associated with the sample point and then divided by (K-1).
[0073] To further illustrate, taking the second sample subset C1 as an example, the average weighted distance of the m-th sample point C1_m in C1 is: dist_C1_m = (cosine distance between C1_m and the centroid of C2 * total number of sample points in C2 + cosine distance between C1_m and the centroid of C3 * total number of sample points in C3 + ... + cosine distance between C1_m and the centroid of Ck-1 * total number of sample points in Ck-1) / (K-1), where C2, C3, and Ck-1 are the other K-1 second sample subsets. The first cosine distance mentioned above includes the cosine distance between C1_m and the centroid of C2, and the second cosine distance mentioned above includes the cosine distance between C1_m and the centroid of C2 * total number of sample points in C2.
[0074] To further illustrate, such as Figure 3 As shown (where black spheres represent centroid points and white spheres represent sample points that are not centroid points), the first sample subset 302 includes four sample points. Sample point 3002 is the centroid point, and sample points 3004 and 3006 are sample points of the second sample subset 304 after edge selection processing corresponding to the first sample subset 302. The average weighted distance dist1 of sample point 3004 is (d1*1+d2*1) / 2; the average weighted distance dist2 of sample point 3006 is (d3*1+d4*1) / 2. Furthermore, when dist1 is greater than dist2, sample point 3004 is determined as the representative sample point of the second sample subset 304; and when dist2 is greater than dist1, sample point 3006 is determined as the representative sample point of the second sample subset 304.
[0075] Through the embodiments provided in this application, the current i-th second sample subset is determined from K second sample subsets, and the average weighted distance between each sample point in the i-th second sample subset and K-1 initial centroids is obtained; wherein, obtaining the average weighted distance between the m-th sample point in the i-th second sample subset and K-1 initial centroids includes: obtaining the cosine distance between the m-th sample point and each centroid of the K-1 initial centroids respectively, to obtain K-1 first cosine distance values; obtaining each cosine distance value in the K-1 cosine distance values. The product of the chord distance and the number of sample points in the second sample subset containing the associated centroid is used to obtain K-1 second cosine distance values. The third cosine distance value obtained by summing the K-1 second cosine distance values is used as the quotient of the third cosine distance value and K-1 to determine the average weighted distance between the m-th sample point and the K-1 initial centroids. The j-th sample point included in the i-th second sample subset is determined as the representative sample point associated with the i-th sample subset, where the average weighted distance of the j-th sample point is the largest among all sample points included in the i-th second sample subset. Based on the initial centroids and the corresponding initial clustering results, and based on the edge selection and weighting of the initial clustering results, new centroids are obtained. The new centroids obtained based on edge selection and weighting are "uniformly distributed" with the initial centroids. Through continuous iteration, a final number of uniformly distributed centroids are obtained, solving the problem of overly close or scattered distribution caused by random selection of centroids, ensuring the uniformity of centroid selection, and thus achieving the technical effect of improving the accuracy of clustering.
[0076] As an alternative approach, after determining the (K+1)th centroid, the method also includes:
[0077] S1, based on K+1 centroids, cluster the other sample points in the sample set to obtain K+1 third sample subsets, where the sample points in each third sample subset are sorted from near to far according to the cosine distance between them and the corresponding centroid.
[0078] S2, perform edge selection processing on K+1 third sample subsets respectively to obtain K+1 fourth sample subsets, wherein each fourth sample subset includes multiple sample points with preset weights and the furthest sorting from the corresponding third sample subset;
[0079] S3, based on all sample points included in each fourth sample subset and the centroids associated with the other K fourth sample subsets, determine the K+2th centroid, where the obtained K+2 centroids are associated with K+2 third sample subsets;
[0080] S4. Repeat the above steps until the preset convergence condition is met.
[0081] Optionally, in this embodiment, after determining the (K+1)th centroid point, if the preset convergence condition is not met, the currently obtained (K+1) centroid points can be determined as the (K+1)th new initial centroid points of the sample set, and clustering can be performed based on the (K+1)th new initial centroid points to obtain the clustered (K+1)th third sample subsets centered on each of the new initial centroid points. Edge selection processing can be performed on the (K+1)th third sample subsets to obtain the corresponding (K+1)th fourth sample subsets. Based on all the sample points included in each third sample subset and the centroid points associated with the other (K-1)th fourth sample subsets, the (K+2)th centroid point is determined, and so on, until the preset convergence condition is met.
[0082] Optionally, in this embodiment, the preset convergence condition may include, but is not limited to, the determination of centroid points reaching a preset number threshold, and may also include, but is not limited to, the cosine distance between each sample point of each obtained second sample subset and the corresponding centroid point being less than or equal to the preset cosine distance.
[0083] It should be noted that, under the condition of satisfying the preset convergence condition, based on the determined centroid, the other sample points in the sample set other than the determined centroid are clustered to obtain the clustered target sample set.
[0084] According to the embodiments provided in this application, based on K+1 centroids, other sample points in the sample set are clustered to obtain K+1 third sample subsets. The sample points in each third sample subset are sorted from closest to farthest according to their cosine distance to the corresponding centroid. Edge selection is then performed on each of the K+1 third sample subsets to obtain K+1 fourth sample subsets. Each fourth sample subset includes multiple sample points from the corresponding third sample subset, with preset weights and the furthest sorted points. Based on all the sample points included in each fourth sample subset and the centroids associated with the other K fourth sample subsets, the K+2th centroid is determined. The obtained K+2 centroids are associated with K+2 third sample subsets. The above steps are repeated until a preset convergence condition is met.
[0085] As an optional approach, the preset convergence conditions include:
[0086] S1, if the number of determined centroids reaches a preset threshold, then the preset convergence condition is satisfied; or
[0087] S2, if the cosine distance between each sample point in each of the obtained second sample subsets and the corresponding centroid point is not greater than the preset cosine distance, then the preset convergence condition is satisfied.
[0088] Optionally, in this embodiment, the preset convergence condition may include, but is not limited to, the determination of centroid points reaching a preset number threshold, and may also include, but is not limited to, the cosine distance between each sample point of each obtained second sample subset and the corresponding centroid point being less than or equal to the preset cosine distance.
[0089] It should be noted that, under the condition of satisfying the preset convergence condition, based on the determined centroid, the other sample points in the sample set other than the determined centroid are clustered to obtain the clustered target sample set.
[0090] Through the embodiments provided in this application, when the number of determined centroids reaches a preset number threshold, it is determined that the preset convergence condition is met; or when the cosine distance between each sample point in each of the obtained second sample subsets and the corresponding centroid is not greater than the preset cosine distance, it is determined that the preset convergence condition is met.
[0091] As an optional approach, obtaining K sample points from a sample set containing N sample points and determining these K sample points as K initial centroids includes:
[0092] S1, iterate through the N sample points to obtain the cosine distances between each pair of samples;
[0093] S2, select at least two sample points with the largest cosine distance, and determine at least two sample points as K sample points.
[0094] Optionally, in this embodiment, obtaining the initial centroid from the sample set may include, but is not limited to, obtaining the cosine distance between each pair of sample points in the sample set, and determining at least two sample points with the largest cosine distance as K sample points, i.e., K initial centroids.
[0095] Through the embodiments provided in this application, the cosine distances between each pair of N sample points are obtained; at least two sample points with the largest cosine distance are selected, and these at least two sample points are determined as K sample points.
[0096] As an alternative approach, before obtaining K sample points from a sample set containing N sample points, the method further includes:
[0097] S1. Obtain a sample set from the historical error log set. The sample set includes error log data of different error categories. The historical error log set includes error log data generated by the account associated with the client during the historical time period.
[0098] After obtaining the clustered target sample set, the method also includes:
[0099] S2, perform feature extraction processing on the target sample set to obtain the erroneous operation behavior features of the account, where the erroneous operation behavior features are used to indicate the account's erroneous operation habits in a historical time period.
[0100] Optionally, in this embodiment, the sample set may be obtained from, but is not limited to, a historical error log set, wherein the historical error log set includes error log data generated by the client-associated account during historical time periods, and may include, but is not limited to, a large amount of error log data of different error categories.
[0101] It should be noted that the clustering method of the sample points in this application is also applicable to log data or sample sets that indicate the usage habits of accounts associated with clients over historical time periods, and this application does not impose specific restrictions on this.
[0102] Optionally, in this embodiment, after obtaining the clustered target sample set, the target sample set is used to analyze the characteristics of erroneous habits in historical time periods.
[0103] It should be noted that by using the above-mentioned clustering method for the sample points to obtain high-quality centroids from the large amount of input historical error log data, multiple log groups can be quickly and effectively obtained. By sorting out the error logs of each group, the user account's incorrect usage habits and corresponding solutions during the use of related applications can be obtained, forming solution assets that provide reference ideas for subsequent operation and maintenance personnel when analyzing similar problems.
[0104] As an alternative approach, the clustering method for the above sample points can be applied to a k-means log clustering scenario based on edge selection and weighted distance. In this scenario, log clustering aims to identify similar logs. By analyzing the errors encountered by users during application usage, logs with the same error type are grouped together. Then, the system can be further categorized to uncover potential errors in user habits and provide suggestions for addressing similar issues in the future.
[0105] Furthermore, in this scenario, text clustering aims to find similar texts, which is highly valuable for data mining. Traditional k-means clustering algorithms randomly select initial centroids during the initialization phase. This can lead to centroids that are either too scattered or too concentrated, resulting in uneven clustering and slow convergence, ultimately leading to poor clustering results. This deficiency is particularly pronounced with massive amounts of text data. Additionally, randomly selecting discrete points or noisy data as initial centroids also hinders subsequent clustering convergence, resulting in unsatisfactory clustering performance.
[0106] Based on the above-mentioned clustering methods for sample points, and addressing the problems of slow convergence and unsatisfactory clustering results when performing K-means clustering on a large number of texts, this embodiment proposes a k-means log classification method based on edge-selected weighted distance, including:
[0107] Step S402: For the log sample set dataset, calculate the cosine distance between each pair of samples, and find the two samples with the greatest distance, denoted as C1 and C2. Divide the samples in the dataset into two clusters according to the principle of being closest to C1 and C2 in cosine distance.
[0108] It should be noted that for cluster C1, the total number of samples in the cluster is N1. The samples in cluster C1 are sorted from closest to farthest from the cosine distance of C1. The samples that are taken out are denoted as the edge set C1_out of cluster C1, and the number of samples in the edge set C1_out is 10%*N1.
[0109] It should be noted that for cluster C2, the total number of samples in the cluster is N2. The samples in cluster C2 are sorted from the nearest to the farthest from the cosine distance of C2. The samples that are taken out are denoted as the marginal set C2_out of cluster C2, and the number of samples in the marginal set C2_out is 10%*N2.
[0110] Step S404: Calculate the weighted distance between samples in the edge set C1_out and C2, and find the sample C1_out_max with the largest weighted distance, as follows:
[0111] The weighted distance between sample x and C2 is the cosine distance between sample x and C2 multiplied by the total number of samples in cluster C2, N2. Calculate the weighted distance between each sample in C1_out and C2, and find the sample with the largest weighted distance, C1_out_max. The corresponding weighted distance is denoted as dist_C1_out.
[0112] Step S406: Calculate the weighted distance between samples in the edge set C2_out and C1, and find the sample C1_out_max with the largest weighted distance, as follows:
[0113] The weighted distance between sample x and C1 is the cosine distance between sample x and C1 multiplied by the total number of samples in cluster C1, N1. Calculate the weighted distance between each sample in C2_out and C1, and find the sample with the largest weighted distance, C2_out_max. The corresponding weighted distance is denoted as dist_C2_out.
[0114] Step S408: Compare the weighted distances dist_C1_out and dist_C2_out to obtain the larger weighted distance. The sample corresponding to the larger weighted distance is denoted as C3. For example, if dist_C1_out > dist_C2_out, then the sample C1_out_max corresponding to the weighted distance dist_C1_out is denoted as C3.
[0115] Step S410: Determine whether the number of centroids that have been determined has reached K.
[0116] Step S412: If the number of centroids is less than K, divide the sample set into clusters of a total number of centroids according to the principle of the closest cosine distance to the selected centroids.
[0117] It should be noted that when the number of centroids reaches K, the initialization of centroids ends, and subsequent clustering is performed based on the determined centroids.
[0118] Step S414: Find the edge set of each cluster, calculate the average weighted distance between the edge sample of each cluster and the other selected cluster centers, and find the sample with the largest average weighted distance and the largest weighted distance of the cluster.
[0119] Step S416: Compare the maximum average weighted distance of all clusters, obtain the maximum average weighted distance, record the sample corresponding to the average weighted distance as the next centroid, and return to step S410 to determine whether the number of centroids currently determined has reached K.
[0120] To further illustrate, based on the above step S408, after obtaining three centroids C1, C2, and C3 and their corresponding clusters, and with K greater than 3, the edge sets of the three clusters are calculated respectively. That is, the samples within the cluster are sorted from near to far by calculating the cosine distance between the samples and the cluster center. After taking them out, 90% to 100% of the samples are recorded as the edge set of the cluster. The number of edge set samples is 10% of the number of samples in the corresponding cluster.
[0121] Calculate the average weighted distance between samples in the edge set of cluster C1 and the other cluster centers (i.e., C2, C3), and find the sample C1_out_max corresponding to the maximum average weighted distance, as follows:
[0122] The average weighted distance between sample x, the edge set of cluster C1, and the other cluster centers (i.e., C2 and C3) is calculated as (cosine distance between sample x and C2 * total number of samples in cluster C2 + cosine distance between sample x and C3 * total number of samples in cluster C3) / total number of other cluster centers. In this case, the total number of other cluster centers is 2. Find the sample C1_out_max with the largest average weighted distance to the other cluster centers (i.e., C2 and C3), and denote the corresponding maximum weighted distance as dist_C1_out.
[0123] Calculate the average weighted distance between samples in the edge set of cluster C2 and the other cluster centers (i.e., C1, C3), and find the sample C2_out_max corresponding to the maximum average weighted distance, as follows:
[0124] The average weighted distance between sample x, the edge set of cluster C2, and the other cluster centers (i.e., C1 and C3) is calculated as (cosine distance between sample x and C1 * total number of samples in cluster C1 + cosine distance between sample x and C3 * total number of samples in cluster C3) / total number of other cluster centers. In this case, the total number of other cluster centers is 2. Find the sample C2_out_max with the largest average weighted distance to the other cluster centers (i.e., C1 and C3), and denote the corresponding maximum weighted distance as dist_C2_out.
[0125] Calculate the average weighted distance between samples in the edge set of cluster C3 and the other cluster centers (i.e., C1, C2), and find the sample C3_out_max corresponding to the maximum average weighted distance, as follows:
[0126] The average weighted distance between sample x, the edge set of cluster C3, and the other cluster centers (i.e., C1 and C2) is calculated as (cosine distance between sample x and C1 * total number of samples in cluster C1 + cosine distance between sample x and C2 * total number of samples in cluster C2) / total number of other cluster centers. In this case, the total number of other cluster centers is 2. Find the sample C3_out_max with the largest average weighted distance to the other cluster centers (i.e., C1 and C2), and denote the corresponding maximum weighted distance as dist_C3_out.
[0127] Compare the average weighted distances dist_C1_out, dist_C2_out, and dist_C3_out to find the largest weighted distance. The sample corresponding to this weighted distance is denoted as C4. For example, if dist_C1_out is greater than dist_C2_out and dist_C3_out, then the sample C1_out_max corresponding to the average weighted distance dist_C1_out is denoted as C4.
[0128] Similarly...
[0129] The dataset is divided into k-1 clusters based on the principle of being closest to the cosine distance of C1, C2, C3, ..., Ck-1.
[0130] Find the edge set of each cluster, calculate the maximum average weighted distance between the edge set of each cluster and the remaining selected cluster centers. Each cluster has a maximum average weighted distance. Select the largest value from these maximum average weighted distances; the sample corresponding to this value is Ck. The details are as follows:
[0131] Find the edge sample set of cluster C1 (i.e., the 10% of samples in this cluster that are farthest from cluster C1 by cosine distance). Calculate the average weighted distance between the samples in the edge set of cluster C1 and the remaining cluster centers (i.e., C2, C3, ..., Ck-1). Find the sample C1_out_max corresponding to the maximum average weighted distance. Details:
[0132] The average weighted distance between sample x, the edge set of cluster C1, and the other cluster centers (C2, C3, ..., Ck-1) is calculated as: (cosine distance between sample x and C2 * total number of samples in cluster C2 + cosine distance between sample x and C3 * total number of samples in cluster C3 + ... + cosine distance between sample x and Ck-1 * total number of samples in cluster Ck-1) / total number of other cluster centers, where the total number of other cluster centers is k-1. The sample C1_out_max with the largest average weighted distance to the other cluster centers (C2, C3, ..., Ck-1) is then identified, and the corresponding maximum weighted distance is denoted as dist_C1_out. ......
[0134] Find the edge sample set of cluster Ck-1 (i.e., the 10% of samples in this cluster that are farthest from cluster Ck-1 by cosine distance). Calculate the average weighted distance between the samples in the edge set of cluster Ck-1 and the remaining cluster centers (i.e., C1, C2, C3, ..., Ck-2). Find the sample Ck-1_out_max corresponding to the maximum average weighted distance. Details:
[0135] The average weighted distance between a sample x, an edge set of cluster Ck-1, and the other cluster centers (C1, C2, C3, ..., Ck-2) is calculated as: (cosine distance between sample x and C1 * total number of samples in cluster C1 + cosine distance between sample x and C2 * total number of samples in cluster C2 + ... + cosine distance between sample x and Ck-2 * total number of samples in cluster Ck-2) / total number of the other cluster centers, where the total number of the other cluster centers is k-1. The sample Ck-1_out_max with the largest average weighted distance to the other cluster centers (C1, C2, C3, ..., Ck-2) is then identified, and this maximum weighted distance is denoted as dist_Ck-1_out.
[0136] Find the largest value among dist_C1_out, dist_C2_out, ..., dist_Ck-1_out, and denote the sample corresponding to this value as Ck.
[0137] At this point, C1, C2, ..., Ck are the initial centroids, and k-means clustering is performed.
[0138] It should be noted that the initialization phase of the kmeans algorithm was improved using the above method. A weighted distance method based on edge selection was proposed to disperse the centroid selection during the initialization process. This accelerates the initialization convergence speed and improves the quality of initial centroid selection. The clustering quality is better, resulting in more accurate classification of log data.
[0139] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0140] This application also provides a clustering device for sample points. It should be noted that the clustering device for sample points in this application can be used to execute the clustering method for sample points provided in this application. The clustering device for sample points provided in this application will be described below.
[0141] Figure 5 This is a schematic diagram of a clustering device for sample points according to an embodiment of this application. Figure 5 As shown, the device includes:
[0142] The acquisition unit 502 is used to acquire K sample points from a sample set including N sample points, and determine the K sample points as K initial centroid points, where N is a positive integer greater than 2 and K is a positive integer less than N;
[0143] The first clustering unit 504 is used to cluster other sample points in the sample set based on K initial centroids to obtain K first sample subsets centered on each initial centroid. The sample points in each first sample subset are sorted from near to far according to the cosine distance between them and the corresponding initial centroid.
[0144] Edge selection unit 506 is used to perform edge selection processing on K first sample subsets respectively to obtain K second sample subsets, wherein each second sample subset includes multiple sample points with preset weights and the furthest order within the corresponding first sample subset;
[0145] The determining unit 508 is used to determine the K+1th centroid point based on all sample points included in each second sample subset and each centroid point associated with the other K-1 second sample subsets, wherein the obtained K+1 centroid points are associated with K+1 first sample subsets;
[0146] The second clustering unit 510 is used to cluster other sample points in the sample set, excluding the determined centroid points, based on the determined centroid points, under the condition of satisfying the preset convergence conditions, so as to obtain the clustered target sample set.
[0147] Optionally, in the clustering apparatus for sample points provided in this application embodiment, the determining unit 508 includes:
[0148] The first determining module is used to determine the representative sample point associated with each second sample subset based on the average weighted distance of all sample points included in each second sample subset to the K-1 initial centroid points associated with the other K-1 second sample subsets, wherein the representative sample point is the sample point with the largest average weighted distance among all second sample subsets.
[0149] The second determination module is used to determine the target sample point with the largest average weighted distance from all K representative sample points, and to determine the target sample point as an additional centroid point, thus obtaining K+1 centroid points.
[0150] Optionally, in the clustering device for sample points provided in the embodiments of this application, the first determining module includes:
[0151] The first determining submodule is used to determine the current i-th second sample subset from the K second sample subsets, and to obtain the average weighted distance between each sample point and the K-1 initial centroid points in the i-th second sample subset.
[0152] The average weighted distance between the m-th sample point in the i-th second sample subset and the K-1 initial centroid points includes:
[0153] The first acquisition submodule is used to acquire the cosine distances between the m-th sample point and each of the K-1 initial centroid points, and obtain the K-1 first cosine distance values;
[0154] The second acquisition submodule is used to obtain the product of each cosine distance in the K-1 cosine distance values and the number of sample points in the second sample subset where the associated centroid point is located, to obtain K-1 second cosine distance values.
[0155] The second determination submodule is used to determine the average weighted distance between the m-th sample point and the K-1 initial centroid points by summing the K-1 second cosine distance values and the third cosine distance value obtained by summing the K-1 second cosine distance values.
[0156] The third determination submodule is used to determine the representative sample point associated with the j-th sample point included in the i-th second sample subset, wherein the average weighted distance of the j-th sample point is the largest among all sample points included in the i-th second sample subset.
[0157] Optionally, in the sample point clustering apparatus provided in the embodiments of this application, the apparatus further includes:
[0158] The clustering module is used to cluster other sample points in the sample set after determining the K+1th centroid point, to obtain K+1 third sample subsets. The sample points in each third sample subset are sorted from near to far according to the cosine distance between them and the corresponding centroid point.
[0159] The edge selection module is used to perform edge selection processing on the K+1 third sample subsets after determining the K+1 centroid point to obtain K+1 fourth sample subsets. Each fourth sample subset includes multiple sample points with preset weights and the furthest order within the corresponding third sample subset.
[0160] The third determining module is used to determine the K+2 centroid point after determining the K+1 centroid point, based on all sample points included in each fourth sample subset and the centroid points associated with the other K fourth sample subsets. The obtained K+2 centroid points are associated with K+2 third sample subsets.
[0161] The execution module is used to repeat the above steps after the K+1th centroid is determined until the preset convergence condition is met.
[0162] Optionally, in the sample point clustering apparatus provided in the embodiments of this application, the apparatus further includes:
[0163] The fourth determining module is used to determine whether a preset convergence condition is met when the number of determined centroids reaches a preset threshold; or
[0164] The fifth determining module is used to determine whether the preset convergence condition is met when the cosine distance between each sample point and its corresponding centroid in each of the obtained second sample subsets is not greater than the preset cosine distance.
[0165] Optionally, in the clustering apparatus for sample points provided in this application embodiment, the acquisition unit 502 includes:
[0166] The first acquisition module is used to iterate and obtain the cosine distance between each pair of N sample points;
[0167] The sixth determination module is used to select at least two sample points with the largest cosine distance and determine at least two sample points as K sample points.
[0168] Optionally, in the sample point clustering apparatus provided in the embodiments of this application, the apparatus further includes:
[0169] Before obtaining K sample points from the sample set containing N sample points, the second acquisition module obtains a sample set from the historical error log set. The sample set includes error log data of different error categories, and the historical error log set includes error log data generated by the client's associated account during historical time periods.
[0170] The device also includes:
[0171] The feature extraction module is used to perform feature extraction on the target sample set after obtaining the clustered target sample set to obtain the erroneous operation behavior features of the account. The erroneous operation behavior features are used to indicate the account's erroneous operation habits in a historical time period.
[0172] The sample point clustering device provided in this application embodiment obtains new centroids based on initial centroids and corresponding initial clustering results, and by edge selection and weighting processing based on the initial clustering results. The new centroids obtained based on edge selection and weighting processing have a "uniform distribution relationship" with the initial centroids. Through continuous iteration, a final number of uniformly distributed centroids are obtained, which solves the problem of overly close or scattered distribution caused by random selection of centroids, ensures the uniformity of centroid selection, and thus achieves the technical effect of improving the accuracy of clustering.
[0173] The clustering device for the sample points includes a processor and a memory. The aforementioned acquisition unit, first clustering unit, edge selection unit, determination unit, second clustering unit, etc., are all stored in the memory as program units. The processor executes the aforementioned program units stored in the memory to realize the corresponding functions.
[0174] The processor contains a kernel, which retrieves the corresponding program units from memory. One or more kernels can be configured, and adjusting kernel parameters can ensure the uniformity of centroid selection, thereby improving clustering accuracy.
[0175] The memory may include non-permanent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0176] This invention provides a computer-readable storage medium storing a program that, when executed by a processor, implements a clustering method for the sample points.
[0177] This invention provides a processor for running a program, wherein the program executes a clustering method for the sample points during runtime.
[0178] like Figure 6As shown, an embodiment of the present invention provides an electronic device, which includes a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, it performs the following steps:
[0179] From a sample set containing N sample points, obtain K sample points and determine the K sample points as K initial centroids, where N is a positive integer greater than 2 and K is a positive integer less than N;
[0180] Based on K initial centroids, the other sample points in the sample set are clustered to obtain K first sample subsets centered on each initial centroid. The sample points in each first sample subset are sorted from near to far according to the cosine distance between them and the corresponding initial centroid.
[0181] Edge selection is performed on each of the K first sample subsets to obtain K second sample subsets. Each second sample subset includes multiple sample points with preset weights and the furthest order from the corresponding first sample subset.
[0182] Based on all the sample points included in each second sample subset and the centroids associated with the other K-1 second sample subsets, the K+1th centroid is determined, wherein the obtained K+1 centroids are associated with K+1 first sample subsets;
[0183] If the preset convergence condition is met, based on the determined centroid, the other sample points in the sample set are clustered to obtain the clustered target sample set.
[0184] As an optional approach, based on all sample points included in each second sample subset and the centroids associated with the other K-1 second sample subsets, the K+1th centroid is determined to include:
[0185] Based on the average weighted distance of all sample points included in each second sample subset to the K-1 initial centroids associated with the other K-1 second sample subsets, the representative sample point associated with each second sample subset is determined, where the representative sample point is the sample point with the largest average weighted distance among all second sample subsets.
[0186] From all K representative sample points, the target sample point with the largest average weighted distance is determined, and the target sample point is determined as the additional centroid point, resulting in K+1 centroid points.
[0187] As an optional approach, the representative sample points associated with each second sample subset are determined as follows:
[0188] Determine the current i-th second sample subset from the K second sample subsets, and obtain the average weighted distance between each sample point and the K-1 initial centroids in the i-th second sample subset;
[0189] The average weighted distance between the m-th sample point in the i-th second sample subset and the K-1 initial centroid points includes:
[0190] Obtain the cosine distances between the m-th sample point and each of the K-1 initial centroids to obtain the K-1 first cosine distance values;
[0191] The product of each cosine distance in the K-1 cosine distance values and the number of sample points in the second sample subset where the associated centroid point is located is obtained to get the K-1 second cosine distance values;
[0192] The quotient of the third cosine distance obtained by summing the K-1 second cosine distance values and K-1 is determined as the average weighted distance between the m-th sample point and the K-1 initial centroid points;
[0193] The j-th sample point included in the i-th second sample subset is used to determine the representative sample point associated with the i-th sample subset, where the average weighted distance of the j-th sample point is the largest among all sample points included in the i-th second sample subset.
[0194] As an alternative approach, after determining the (K+1)th centroid, the method also includes:
[0195] Based on K+1 centroids, the other sample points in the sample set are clustered to obtain K+1 third sample subsets. The sample points in each third sample subset are sorted from near to far according to the cosine distance between them and the corresponding centroid.
[0196] Edge selection is performed on each of the K+1 third sample subsets to obtain K+1 fourth sample subsets. Each fourth sample subset includes multiple sample points from the corresponding third sample subset, with preset weights and the furthest order.
[0197] Based on all sample points included in each fourth sample subset and the centroids associated with the other K fourth sample subsets, the K+2th centroid is determined, wherein the K+2 centroids are associated with K+2 third sample subsets.
[0198] Repeat the above steps until the preset convergence condition is met.
[0199] As an optional approach, the preset convergence conditions include:
[0200] If the number of determined centroids reaches a preset threshold, the preset convergence condition is satisfied; or
[0201] If the cosine distance between each sample point in each of the obtained second sample subsets and its corresponding centroid is not greater than the preset cosine distance, then the preset convergence condition is satisfied.
[0202] As an optional approach, obtaining K sample points from a sample set containing N sample points and determining these K sample points as K initial centroids includes:
[0203] Iterate through the N sample points to obtain the cosine distances between each pair of samples;
[0204] Select at least two sample points with the largest cosine distance, and determine these at least two sample points as K sample points.
[0205] As an alternative approach, before obtaining K sample points from a sample set containing N sample points, the method further includes:
[0206] Obtain a sample set from the historical error log collection. The sample set includes error log data of different error categories. The historical error log collection includes error log data generated by the client's associated account during historical time periods.
[0207] After obtaining the clustered target sample set, the method also includes:
[0208] Feature extraction is performed on the target sample set to obtain the erroneous operation behavior features of the account. These erroneous operation behavior features are used to indicate the account's erroneous operation habits over a historical period.
[0209] The devices mentioned in this article can be servers, PCs, tablets, mobile phones, etc.
[0210] This application also provides a computer program product, which, when executed on a data processing device, is suitable for executing a program that initializes the following method steps:
[0211] From a sample set containing N sample points, obtain K sample points and determine the K sample points as K initial centroids, where N is a positive integer greater than 2 and K is a positive integer less than N;
[0212] Based on K initial centroids, the other sample points in the sample set are clustered to obtain K first sample subsets centered on each initial centroid. The sample points in each first sample subset are sorted from near to far according to the cosine distance between them and the corresponding initial centroid.
[0213] Edge selection is performed on each of the K first sample subsets to obtain K second sample subsets. Each second sample subset includes multiple sample points with preset weights and the furthest order from the corresponding first sample subset.
[0214] Based on all the sample points included in each second sample subset and the centroids associated with the other K-1 second sample subsets, the K+1th centroid is determined, wherein the obtained K+1 centroids are associated with K+1 first sample subsets;
[0215] If the preset convergence condition is met, based on the determined centroid, the other sample points in the sample set are clustered to obtain the clustered target sample set.
[0216] As an optional approach, based on all sample points included in each second sample subset and the centroids associated with the other K-1 second sample subsets, the K+1th centroid is determined to include:
[0217] Based on the average weighted distance of all sample points included in each second sample subset to the K-1 initial centroids associated with the other K-1 second sample subsets, the representative sample point associated with each second sample subset is determined, where the representative sample point is the sample point with the largest average weighted distance among all second sample subsets.
[0218] From all K representative sample points, the target sample point with the largest average weighted distance is determined, and the target sample point is determined as the additional centroid point, resulting in K+1 centroid points.
[0219] As an optional approach, the representative sample points associated with each second sample subset are determined as follows:
[0220] Determine the current i-th second sample subset from the K second sample subsets, and obtain the average weighted distance between each sample point and the K-1 initial centroids in the i-th second sample subset;
[0221] The average weighted distance between the m-th sample point in the i-th second sample subset and the K-1 initial centroid points includes:
[0222] Obtain the cosine distances between the m-th sample point and each of the K-1 initial centroids to obtain the K-1 first cosine distance values;
[0223] The product of each cosine distance in the K-1 cosine distance values and the number of sample points in the second sample subset where the associated centroid point is located is obtained to get the K-1 second cosine distance values;
[0224] The quotient of the third cosine distance obtained by summing the K-1 second cosine distance values and K-1 is determined as the average weighted distance between the m-th sample point and the K-1 initial centroid points;
[0225] The j-th sample point included in the i-th second sample subset is used to determine the representative sample point associated with the i-th sample subset, where the average weighted distance of the j-th sample point is the largest among all sample points included in the i-th second sample subset.
[0226] As an alternative approach, after determining the (K+1)th centroid, the method also includes:
[0227] Based on K+1 centroids, the other sample points in the sample set are clustered to obtain K+1 third sample subsets. The sample points in each third sample subset are sorted from near to far according to the cosine distance between them and the corresponding centroid.
[0228] Edge selection is performed on each of the K+1 third sample subsets to obtain K+1 fourth sample subsets. Each fourth sample subset includes multiple sample points from the corresponding third sample subset, with preset weights and the furthest order.
[0229] Based on all sample points included in each fourth sample subset and the centroids associated with the other K fourth sample subsets, the K+2th centroid is determined, wherein the K+2 centroids are associated with K+2 third sample subsets.
[0230] Repeat the above steps until the preset convergence condition is met.
[0231] As an optional approach, the preset convergence conditions include:
[0232] If the number of determined centroids reaches a preset threshold, the preset convergence condition is satisfied; or
[0233] If the cosine distance between each sample point in each of the obtained second sample subsets and its corresponding centroid is not greater than the preset cosine distance, then the preset convergence condition is satisfied.
[0234] As an optional approach, obtaining K sample points from a sample set containing N sample points and determining these K sample points as K initial centroids includes:
[0235] Iterate through the N sample points to obtain the cosine distances between each pair of samples;
[0236] Select at least two sample points with the largest cosine distance, and determine these at least two sample points as K sample points.
[0237] As an alternative approach, before obtaining K sample points from a sample set containing N sample points, the method further includes:
[0238] Obtain a sample set from the historical error log collection. The sample set includes error log data of different error categories. The historical error log collection includes error log data generated by the client's associated account during historical time periods.
[0239] After obtaining the clustered target sample set, the method also includes:
[0240] Feature extraction is performed on the target sample set to obtain the erroneous operation behavior features of the account. These erroneous operation behavior features are used to indicate the account's erroneous operation habits over a historical period.
[0241] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0242] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0243] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0244] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0245] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0246] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0247] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0248] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0249] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0250] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A clustering method for sample points, characterized in that, include: Obtain a sample set including N sample points from the historical error log set, wherein the sample set includes error log data of different error categories, and the historical error log set includes error log data generated by the client-associated account during historical time periods, where N is a positive integer greater than 2; K sample points are obtained from the sample set, and the K sample points are determined as K initial centroids, where K is a positive integer less than N; Based on the K initial centroids, the other sample points in the sample set are clustered to obtain K first sample subsets centered on each initial centroid. The sample points in each first sample subset are sorted from near to far according to the cosine distance between them and the corresponding initial centroid. Edge selection processing is performed on the K first sample subsets respectively to obtain K second sample subsets, wherein each second sample subset includes multiple sample points with preset weights and the furthest sorted from the corresponding first sample subset; Based on all the sample points included in each second sample subset and the centroids associated with the other K-1 second sample subsets, the K+1th centroid is determined; Repeat the above steps until the number of centroids that have been determined reaches the preset number threshold, or the cosine distance between each sample point in each of the obtained second sample subsets and the corresponding centroid is not greater than the preset cosine distance. Based on the determined centroid, the other sample points in the sample set, excluding the determined centroid, are clustered to obtain the clustered target sample set. The target sample set is subjected to feature extraction processing to obtain the erroneous operation behavior features of the account, wherein the erroneous operation behavior features are used to indicate the account's erroneous operation habits in the historical time period; The process of determining the (K+1)th centroid based on all sample points included in each second sample subset and the centroids associated with each of the other K-1 second sample subsets includes: determining representative sample points associated with each second sample subset based on the average weighted distance of all sample points included in each second sample subset and the other K-1 initial centroids associated with each of the other K-1 second sample subsets, wherein the representative sample point is the sample point with the largest average weighted distance among all second sample subsets; determining the target sample point with the largest average weighted distance from all K representative sample points, and determining the target sample point as the additional centroid, thus obtaining the K+1 centroids, wherein the K+1 centroids are uniformly distributed with the K initial centroids.
2. The method according to claim 1, characterized in that, The determination of the representative sample points associated with each second sample subset includes: Determine the current i-th second sample subset from the K second sample subsets, and obtain the average weighted distance between each sample point and the K-1 initial centroids in the i-th second sample subset; The process of obtaining the average weighted distance between the m-th sample point in the i-th second sample subset and the K-1 initial centroid points includes: Obtain the cosine distances between the m-th sample point and each of the K-1 initial centroids to obtain the K-1 first cosine distance values; The product of each cosine distance in the K-1 cosine distance values and the number of sample points in the second sample subset where the associated centroid point is located is obtained to get K-1 second cosine distance values; The quotient of the third cosine distance obtained by summing the K-1 second cosine distance values and K-1 is determined as the average weighted distance between the m-th sample point and the K-1 initial centroid points; The j-th sample point included in the i-th second sample subset is determined as the representative sample point associated with the i-th second sample subset, wherein the average weighted distance of the j-th sample point is the largest among all sample points included in the i-th second sample subset.
3. The method according to any one of claims 1-2, characterized in that, The step of obtaining K sample points from the sample set and determining the K sample points as K initial centroids includes: Iterate through the N sample points to obtain the cosine distances between each pair of the N sample points; Select at least two sample points with the largest cosine distance, and determine the at least two sample points as the K sample points.
4. The method according to claim 1, characterized in that, After determining the (K+1)th centroid, the method further includes: Based on the K+1 centroids, the other sample points in the sample set are clustered to obtain K+1 third sample subsets. The sample points in each third sample subset are sorted from near to far according to the cosine distance between them and the corresponding centroid. Edge selection processing is performed on the K+1 third sample subsets respectively to obtain K+1 fourth sample subsets, wherein each fourth sample subset includes multiple sample points with preset weights and the furthest sorting from the corresponding third sample subset; Based on all sample points included in each fourth sample subset and the centroids associated with the other K fourth sample subsets, the K+2th centroid is determined, wherein the K+2 centroids are associated with K+2 third sample subsets. Repeat the above steps until the preset convergence condition is met.
5. A clustering device for sample points, characterized in that, include: The device is used to obtain a sample set including N sample points from a historical error log set, wherein the sample set includes error log data of different error categories, and the historical error log set includes error log data generated by the client-associated account during historical time periods, where N is a positive integer greater than 2; The acquisition unit is used to acquire K sample points from the sample set and determine the K sample points as K initial centroids, where K is a positive integer less than N; The first clustering unit is used to cluster other sample points in the sample set based on the K initial centroids to obtain K first sample subsets centered on each initial centroid. The sample points in each first sample subset are sorted from near to far according to the cosine distance between them and the corresponding initial centroid. An edge selection unit is used to perform edge selection processing on the K first sample subsets respectively to obtain K second sample subsets, wherein each second sample subset includes multiple sample points with preset weights and the furthest order within the corresponding first sample subset; The determining unit is used to determine the K+1th centroid based on all sample points included in each second sample subset and the centroids associated with the other K-1 second sample subsets; The device is also used to repeatedly perform the above steps until the number of determined centroids reaches a preset number threshold, or the cosine distance between each sample point in each of the obtained second sample subsets and the corresponding centroid is not greater than the preset cosine distance. The second clustering unit is used to cluster other sample points in the sample set, excluding the determined centroid points, based on the determined centroid points, to obtain the clustered target sample set. The device is further configured to perform feature extraction processing on the target sample set to obtain the erroneous operation behavior features of the account, wherein the erroneous operation behavior features are used to indicate the erroneous operation habits of the account in the historical time period. The apparatus is further configured to determine representative sample points associated with each second sample subset based on the average weighted distance between all sample points included in each second sample subset and the K-1 initial centroids associated with the other K-1 second sample subsets, wherein the representative sample point is the sample point with the largest average weighted distance among all second sample subsets; determine the target sample point with the largest average weighted distance from all K representative sample points, and determine the target sample point as an additional centroid, thereby obtaining the K+1 centroids, wherein the K+1 centroids are uniformly distributed with the K initial centroids.
6. A processor, characterized in that, The processor is used to run a program, wherein the program executes the method according to any one of claims 1 to 4 when it runs.
7. An electronic device, characterized in that, It includes one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method of any one of claims 1 to 4.
Citation Information
Patent Citations
Text clustering method and device, computer equipment and storage medium
CN112989047A
Log analysis method and system and related components
CN114385468A