Industrial internet big data rapid clustering method and system
By constructing a set of representative points and adjusting their positions, combined with the boundary similarity measurement criterion, the problem of insufficient accuracy of traditional clustering algorithms in industrial Internet big data is solved, and efficient and accurate clustering results are achieved.
Patent Information
- Application Number
- CN202510865482.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-10-17
AI Technical Summary
Traditional clustering algorithms find it difficult to accurately capture data distribution of arbitrary shapes in industrial Internet big data, resulting in poor accuracy of clustering results.
A set of representative points is constructed, and the positions of the representative points are adjusted through the adaptive position optimization algorithm of local manifold learning. Cluster analysis is performed based on the boundary similarity measurement criterion, and finally the results are restored to the original data space.
The clustering accuracy and efficiency are improved, and efficient clustering analysis of industrial Internet big data is achieved.
Smart Images

Figure CN120804742A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of data mining, and particularly relates to an industrial internet big data fast clustering method and system. BACKGROUND
[0002] At present, massive industrial data is being generated and accumulated at an unprecedented speed. Industrial internet realizes the digitization, networking and intelligentization of the whole production process by deeply integrating industrial equipment, systems, products, etc. with the internet. These data cover equipment operation monitoring, production process control, quality detection and supply chain, etc. and contain rich industrial knowledge and potential value, which become important factors to drive the improvement of production efficiency, optimization of process flow, stable operation of equipment and upgrading of product quality. In industrial big data analysis, clustering is a key data processing method, which divides data objects into clusters according to similarity principle, realizes data classification and induction, and mines potential patterns and rules to provide solid basis for decision-making, fault diagnosis, predictive maintenance, etc. For example, in equipment fault diagnosis, it can accurately distinguish between normal and fault data to realize early warning and accurate diagnosis of faults.
[0003] However, various complex factors in industrial production process, such as diversity of process flow, difference of equipment performance and dynamic change of environmental conditions, result in high complexity and irregularity of industrial data distribution in data space. Traditional clustering algorithms (such as K-means algorithm) assume that data obeys regular distribution such as spherical or convex shape, and it is difficult to accurately capture the internal structure of data when dealing with arbitrary shape of data distribution in industrial internet big data, resulting in poor accuracy of clustering results. Therefore, how to realize efficient and accurate clustering analysis of industrial internet big data has become a key technical problem to be solved in the current industrial internet field. SUMMARY
[0004] To solve the above technical problems, the application provides an industrial internet big data fast clustering method, which first constructs a representative point set of original data, and then adjusts the position of the representative point by using an adaptive position optimization algorithm based on local manifold learning to enhance its potential clustering structure. On this basis, the optimized representative points are clustered according to the boundary similarity measurement criterion, and finally the clustering results are restored to the original data space by establishing the neighbor relationship between the representative points and the original data.
[0005] To solve the above technical problems, the application provides the following technical solutions: The industrial internet big data fast clustering method comprises the following steps: Step 1: constructing a representative point set of original data set; Step 2: Adjust the positions of the representative points based on the position adjustment matrix to obtain optimized representative points, forming an optimized representative point set; Step 3: Perform cluster analysis on the optimized representative points based on the boundary similarity measurement criterion; Step 4: Through the neighbor relationship between the optimized representative points and the original data, the clustering results are restored to the original data space and the final clustering results are output.
[0006] Furthermore, the step 1 includes: From the original dataset D Random selection m samples as the initial cluster centers; Use clustering algorithm to generate m Clustering results of initial data clusters DR ={ DR 1, …, DR m}; Calculate the i Initial data cluster DRi The cluster center Ci And constitute the representative point set of the original data set DC ={ C 1,…, Cm}; in, DR 1, …, DR m Respectively represent the first,..., m Initial data clusters; C 1,…, Cm Respectively represent the first, ..., m The cluster centers of the initial data clusters.
[0007] Furthermore, the clustering algorithm is a K-means algorithm.
[0008] Furthermore, the step 2 includes: Calculate each representative point Ci The local density LDi ; For each representative point Ci Sure k Neighbor Set KCi ; Constructing correlation weight matrix based on local density and nearest neighbor relationship W ; Adjusting the matrix by position R Adjust the position of the representative point to obtain the optimized representative point , forming the optimized representative point set .
[0009] Furthermore, the local density LDi The calculation formula is:
[0010] where | DR i | represents the initial data cluster DR i The number of data samples included; dist ( D t , C i ) represents a sample D t and representative points C i The Euclidean distance between .
[0011] Furthermore, the correlation weight matrix W The calculation of the elements in the formula satisfies:
[0012] in, is the correlation weight matrix W Middle i Rank j The elements of the column represent the i Representative points Ci With the j Representative points Cj the correlation between LD i and LD j Respectively C i and C j The local density of KC i and KC j Respectively C i and C j of k Neighbor set; dist ( C i , C j )express C i and C j The Euclidean distance of express C i and C kEuclidean distance of denotes C j denotes C k Euclidean distance of
[0013] Further, the position adjustment matrix is:
[0014] wherein, I ∈R m×m denotes an identity matrix, α denotes a given adjustment parameter, denotes a matrix inverse of
[0015] Further, the boundary similarity measure criterion in step 3 is set as:
[0016] wherein, bsim ( CS i , CS j ) denotes the boundary similarity between any two data clusters CS i and CS j in the initialized data cluster set by the optimized representative point set, denotes a set of representative points in CS i which have a near-neighbor relationship with the representative points in CS j ; denotes the number of representative points contained in the set M ij ; denotes a set of representative points in CS j which have a near-neighbor relationship with the representative points in CS i ; denotes the number of representative points contained in the set M ji ; and respectively CS i and CS j denotes the number of representative points contained in the set is the optimized representative point Ci , is the optimized representative point Cj , is The set of k nearest neighbors.
[0017] Furthermore, the step 4 includes: Assign the samples to the cluster to which the nearest optimized representative point belongs; Merge samples with the same label to output the final clustering result.
[0018] In another aspect, the present invention provides an industrial Internet big data rapid clustering system, comprising: Original representative point set construction module: it is used to construct the representative point set of the original data set; Representative point optimization module: used to adjust the positions of the representative points based on the position adjustment matrix to obtain optimized representative points and form an optimized representative point set; Cluster analysis module: It is used to perform cluster analysis on the optimized representative points based on the boundary similarity measurement criterion; Result restoration module: It is used to restore the clustering results to the original data space through the neighbor relationship between the optimized representative points and the original data, and output the final clustering results.
[0019] Compared with the prior art, the present invention has the following beneficial effects: By introducing a set of representative points, this method transforms the clustering problem of large-scale datasets into processing a small number of representative points, reducing the computational scale and overall computational complexity. Furthermore, by optimizing the positions of representative points using a position adjustment matrix, it can more accurately explore the underlying structure and distribution characteristics within the data, effectively improving clustering accuracy and providing a highly efficient solution for cluster analysis in the context of industrial internet big data. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0021] Figure 1 Flowchart of an embodiment of the present invention.
[0022] Figure 2 Figure 1 shows the clustering results of an embodiment of the present invention on the data set t48k. (a) is the original data set, (b) is the representative point set, (c) is the representative point set after position adjustment, (d) is the representative point clustering result, and (e) is the original data set clustering result. DETAILED DESCRIPTION
[0023] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.
[0024] Example 1 The present invention will be further described below with reference to the accompanying drawings.
[0025] In this embodiment, a two-dimensional data set t48k is used for testing. The data set includes 8,000 data samples, the data set dimension is 2, and the data samples are divided into 6 different data categories.
[0026] Figure 1 The flowchart of the embodiment of the present invention is described in the following. The present invention provides a fast clustering method for industrial Internet big data, and defines the industrial Internet data set to be clustered and analyzed as D ={ D 1, …, D n},in D i ={ D i1 ,…, D id} is the first i samples, by d Dimensional features, n is the total number of samples, D ij Corresponding sample D i In the j The values on the dimensions; any two samples D i and D j The Euclidean distance between dist ( D i , D j ) is calculated as .
[0027] Based on the above definition, the method comprises the following steps: Step 1: Construct a set of representative points of the original data set; From the dataset D Random selection m{1, …, 200} are used as initial cluster centers, the K-means clustering algorithm is used to cluster the data set D , and a clustering result is obtained DR = DR 1, …, DR m} is obtained, where the i-th data cluster is represented by DR i . Subsequently, the clustering center i of each data cluster DR i is calculated C i and is used as a representative point of the data set D , and all clustering centers are combined to form a representative point set DC = C 1, …, C m} is obtained, where the i-th clustering center is represented by i C i = C i1 , C i2 , …, C id} is composed of d features, and the value of the j-th feature of the i-th clustering center is represented by d j C ij . The specific calculation method is as follows:
[0028] where DR i represents the number of data samples included in the data cluster DR i ; D tj represents the value of the sample D i in the j-th dimension. Step 2: Calculate the local density of each representative point j DC i in C i , and the specific calculation formula is as follows: LD
[0029] where DR i represents the number of data samples included in the data cluster DR i ; dist (D t , C i ) represents a sample D t and representative points C i The Euclidean distance between DC Each representative point in C i ,calculate C i With each of the remaining representative points Euclidean distance between dist ( C i , C j ), and select the one with the smallest Euclidean distance k (The value is 4) representative points as C i of k Neighbors, put into the set KC i middle.
[0030] Calculate the representative point set DC The correlation weight matrix W ∈R m×m , among which i Rank j Elements of a column express C i and C j The specific calculation formula for the correlation between them is:
[0031] in LD i and LD j Respectively C i and C j The local density of KC i and KC j Respectively C i and C j of k Neighbor set; dist ( C i , C j )express C i andC j The Euclidean distance.
[0032] Calculate a dimension as Position adjustment matrix ,in I ∈R m×m The identity matrix represented by α Represents a given tuning parameter , in this embodiment, the value is 0.). W represents the correlation weight matrix, Representation matrix The inverse matrix of .
[0033] Step 2: Adjust the positions of the representative points based on the position adjustment matrix to obtain optimized representative points, forming an optimized representative point set; right DC Each representative point in C i After position adjustment, the adjusted representative point is ,in express No. j Then all the adjusted representative points are combined into a new .in The specific calculation is:
[0034] in C ij express C i No. j The value of each dimension; R it Representation matrix R No. i Rank t The value of the column.
[0035] Step 7: For each adjustment representative point ,calculate With each of the remaining adjustments representative points Euclidean distance between , and select the 4 adjustment representative points with the smallest Euclidean distance as of k Neighbors, put into the set middle.
[0036] Each adjusted representative point Initialize a data cluster , forming the initial clustering result set of representative points .
[0037] Step 3: Perform cluster analysis on the optimized representative points based on the boundary similarity measurement criterion; calculate CS Any two data clusters CS i and CS j The boundary similarity between bsim ( CS i , CS j ), which is calculated as follows:
[0038] in Indicated by CS i Zhongyu CS j The set of representative points that have a neighboring relationship with the representative points in ; Representing a collection M ij The number of representative points included; Indicated by CS j Zhongyu CS i The set of representative points that have a neighboring relationship with the representative points in ; Representing a collection M ji The number of representative points included; and respectively CS i and CS j The number of representative points included.
[0039] Select CS The two data clusters with the highest similarity in the middle boundary CS i and CS j Merge to get a new data cluster ,from CS Remove data clusters CS i and CS j , and the new data cluster Add to CS .
[0040] If the current CS The number of clusters in has reached the preset threshold τ (value is 6), then CSeach data cluster as a class, and assign a label from 1 to τ Step 4 is executed; if the number of clusters in CS does not reach τ , the boundary similarity calculation step is returned.
[0041] Step 4: The clustering result is restored to the original data space through the neighbor relationship between the optimized representative points and the original data, and the final clustering result is output.
[0042] For each sample D i in the data set, D i , the Euclidean distance between CS and all representative points in each clustering cluster is calculated, and the label of the representative point with the smallest distance is assigned to D i . Finally, data samples with the same label are classified into a class, and the final clustering result is output.
[0043] See Figure 2 , which is the clustering result of the embodiment of the present application on the experimental data set. It can be seen that the present application can effectively perform clustering analysis on the given data set on this embodiment.
[0044] The present application proposes a fast clustering method for industrial internet big data. The method effectively improves the speed and efficiency of clustering by clustering the representative points of the original data samples. At the same time, the position adjustment matrix is used to strengthen the potential clustering structure of the representative points, ensuring the clustering accuracy and realizing efficient clustering analysis of industrial internet big data.
[0045] Embodiment 2 The embodiment provides a fast clustering system for industrial internet big data, comprising: An original representative point set construction module is used to construct a representative point set of the original data set; A representative point optimization module is used to adjust the position of the representative point based on the position adjustment matrix to obtain an optimized representative point, and to form an optimized representative point set; A clustering analysis module is used to perform clustering analysis on the optimized representative point based on the boundary similarity measurement criterion; A result restoration module is used to restore the clustering result to the original data space through the neighbor relationship between the optimized representative point and the original data, and to output the final clustering result.
[0046] It should be understood that the parts not elaborated in the specification are all prior art.
[0047] It should be understood that the above description is merely a detailed example of the preferred embodiment and is not to be taken in a limiting sense. There can be many variations to the embodiments described herein without departing from the spirit of the application. The scope of the application should be determined by a fair reading of the appended claims, along with the full text of the specification.
Claims
1. The rapid clustering method of industrial Internet big data is characterized by: The following steps are involved: Step 1: Construct a set of representative points of the original data set; Step 2: Adjust the positions of the representative points based on the position adjustment matrix to obtain optimized representative points, forming an optimized representative point set; Step 3: Perform cluster analysis on the optimized representative points based on the boundary similarity measurement criterion; Step 4: Through the neighbor relationship between the optimized representative points and the original data, the clustering results are restored to the original data space and the final clustering results are output.
2. The industrial Internet big data rapid clustering method according to claim 1 is characterized in that: The step 1 comprises: From the original dataset D Random selection m samples as the initial cluster centers; Use clustering algorithm to generate m Clustering results of initial data clusters DR ={ DR 1, …, DR m }; Calculate the i Initial data cluster DRi The cluster center Ci And constitute the representative point set of the original data set DC ={ C 1,…, Cm }; in, DR 1, …, DR m Respectively represent the first,..., m Initial data clusters; C 1,…, Cm Respectively represent the first, ..., m The cluster centers of the initial data clusters.
3. The industrial Internet big data rapid clustering method according to claim 1, characterized in that: The clustering algorithm is the K-means algorithm.
4. The industrial Internet big data rapid clustering method according to claim 1, characterized in that: The step 2 includes: Calculate each representative point Ci The local density LDi ; For each representative point Ci Sure k Neighbor Set KC ; Constructing correlation weight matrix based on local density and nearest neighbor relationship W ; Adjusting the matrix by position R Adjust the position of the representative point to obtain the optimized representative point , forming the optimized representative point set .
5. The industrial Internet big data rapid clustering method according to claim 4 is characterized in that: Local density LDi The calculation formula is: Among them, | DR i | represents the initial data cluster DR i The number of data samples included; dist ( D t , C i ) represents a sample D t and representative points C i The Euclidean distance between .
6. The industrial Internet big data rapid clustering method according to claim 1, characterized in that: The relevance weight matrix W The calculation of the elements in the formula satisfies: in, is the correlation weight matrix W Middle i Rank j The elements of the column represent the i Representative points Ci With the j Representative points Cj the correlation between LD i and LD j Respectively C i and C j The local density of KC i and KC j Respectively C i and C j of k Neighbor set; dist ( C i , C j )express C i and C j The Euclidean distance of express C i and C k The Euclidean distance of express C j and C k The Euclidean distance.
7. The industrial Internet big data rapid clustering method according to claim 6, characterized in that: The position adjustment matrix is: in, I ∈R m×m The identity matrix represented by α represents a given tuning parameter, Representation matrix The inverse matrix of .
8. The industrial Internet big data rapid clustering method according to claim 1, characterized in that: The boundary similarity measurement criterion in step 3 is set as: in, bsim ( CS i , CS j ) represents any two data clusters in the data cluster set initialized by the optimized representative point set CS i and CS j The boundary similarity between Indicated by CS i Zhongyu CS j The set of representative points that have a neighboring relationship with the representative points in ; Representing a collection M ij The number of representative points included; Indicated by CS j Zhongyu CS i The set of representative points that have a neighboring relationship with the representative points in ; Representing a collection M ji The number of representative points included; and respectively CS i and CS j The number of representative points included; Representative point Ci The optimized representative points, Representative point Cj The optimized representative points, for The set of k nearest neighbors.
9. The industrial Internet big data rapid clustering method according to claim 1, characterized in that: The step 4 comprises: Assign the samples to the cluster to which the nearest optimized representative point belongs; Merge samples with the same label to output the final clustering result.
10. Industrial Internet Big Data Rapid Clustering System, characterized by: include: Original representative point set construction module: it is used to construct the representative point set of the original data set; Representative point optimization module: used to adjust the positions of the representative points based on the position adjustment matrix to obtain optimized representative points and form an optimized representative point set; Cluster analysis module: It is used to perform cluster analysis on the optimized representative points based on the boundary similarity measurement criterion; Result restoration module: It is used to restore the clustering results to the original data space through the neighbor relationship between the optimized representative points and the original data, and output the final clustering results; The industrial Internet big data rapid clustering system is used to execute the steps in the industrial Internet big data rapid clustering method described in any one of claims 1-9.