K nearest neighbor search method based on cluster structure optimization and triangular inequality strategy
Through the K nearest neighbor search method based on cluster structure optimization and triangular inequality strategies, the problem of inefficiency of traditional methods in large-scale high-dimensional data processing is solved, and fast and accurate nearest neighbor search is achieved to meet the fault diagnosis needs in complex industrial scenarios.
Patent Information
- Application Number
- CN202510325348.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-20
AI Technical Summary
Traditional nearest neighbor search methods consume huge computing resources when processing large-scale high-dimensional equipment data, and the search efficiency decreases, making it difficult to meet the dual requirements of fault diagnosis for real-time and accuracy.
The K nearest neighbor search method based on cluster structure optimization and triangular inequality strategy is adopted. The sample set is preliminarily divided by the k-means++ clustering algorithm, the representative distance is calculated and the unqualified cluster judgment rules are constructed, the unqualified cluster is decomposed into qualified subclusters, and the K nearest neighbor search is performed in combination with the triangular inequality inspection strategy.
It significantly improves search efficiency, reduces K nearest neighbor search time, and can meet the equipment fault diagnosis needs in complex industrial scenarios.
Smart Images

Figure CN120180256A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of nearest neighbor search, and in particular to a K-nearest neighbor search method based on cluster structure optimization and triangle inequality strategy. Background Art
[0002] In equipment fault diagnosis, quickly and accurately retrieving samples similar to the current state of the equipment from the massive historical data of the equipment is of great significance for realizing early warning and accurate positioning of equipment faults. With the explosive growth of the scale of industrial equipment operation data, the nearest neighbor search technology in the scenario of large-scale high-dimensional data sets has become a key support in the field of state warning and fault diagnosis. However, traditional nearest neighbor search methods face severe challenges when dealing with large-scale high-dimensional equipment data: the consumption of computing resources is huge, the search efficiency drops sharply, and it is difficult to meet the dual requirements of real-time performance and accuracy for fault diagnosis. To address this challenge, researchers have proposed various improvement methods, among which the methods based on planar indexing have received extensive attention due to their high efficiency and scalability. These methods divide the equipment state data space into multiple clusters, and utilize the characteristic that the data distributions within the clusters are similar to narrow the search range, thereby improving the efficiency of fault diagnosis. However, due to insufficient attention to the quantity and quality of clusters, and the lack of adaptability to the non-uniform distribution characteristics of equipment state data, it is difficult to meet the fault diagnosis requirements in complex industrial scenarios.
[0003] In recent years, the nearest neighbor search method based on the triangle inequality checking strategy has shown unique advantages. This method utilizes the property that the distance metric satisfies the triangle inequality, and by pre-computing and storing the distances between some data points, quickly excludes candidate points that do not meet the conditions during the search process, thereby reducing unnecessary distance calculations. However, most of the existing technologies of this method rely on a single clustering strategy, are difficult to effectively process complex data distributions, and lack effective processing of some instances that affect the clustering quality. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides a K-nearest neighbor search method based on cluster structure optimization and triangle inequality strategy, aiming to achieve fast nearest neighbor search under the conditions of large-scale high-dimensional data sets and meet the equipment fault diagnosis requirements in complex industrial scenarios.
[0005] The technical solution adopted by the present invention is as follows:
[0006] The present invention provides a K-nearest neighbor search method based on cluster structure optimization and triangle inequality strategy, including:
[0007] S1. Using the k-means++ clustering algorithm to initially divide the sample set to be searched into M clusters with similar characteristics, and obtaining the cluster structure information of the sample set, which includes the cluster center position, cluster radius, number of samples, and cluster labels of each sample belonging to each cluster;
[0008] S2. Calculate the representative distance δ of each sample within each cluster j,i :
[0009]
[0010] where y j,i represents the i-th sample in the j-th cluster, Nj is the number of samples in y j,i , j = 1,..., M; d(y j,i , y j,q ) represents the Euclidean distance between y j,i and y j,q ;
[0011] Based on the representative distance, construct a non - qualified cluster determination rule, and determine all clusters to obtain non - qualified clusters;
[0012] The determination rule includes:
[0013]
[0014] In the above formula, max(δ j,i ) is the maximum value of the representative distances of all samples within the cluster, λ is a boundary adjustment factor, which is an empirically set value or the optimal value obtained through the Aurora optimization algorithm; if the samples within the cluster satisfy the above formula, it is determined as a qualified cluster, otherwise, it is determined as a non - qualified cluster;
[0015] Among them, the boundary adjustment factor λ and the initially divided number of clusters M are empirically set values or the optimal values obtained through the Aurora optimization algorithm;
[0016] S3. Based on the k - means++ clustering algorithm and the non - qualified cluster determination rule, decompose each non - qualified cluster into multiple qualified sub - clusters, update M and the cluster structure information, and obtain a sample set with updated structure;
[0017] S4. Based on the training sample set with updated structure, use the triangle inequality checking strategy to perform K - nearest neighbor search on the sample to be queried.
[0018] A further technical solution is:
[0019] In step S1, using the k - means++ clustering algorithm to initially divide the sample set to be searched into M clusters with similar characteristics and obtain the cluster structure information of the sample set, including:
[0020] S11. For a sample set T composed of n samples y ∈ R D , randomly select a sample in T as the initial cluster center c1;
[0021] S12. For non - cluster - center samples y* , calculate the shortest distance D(y j ) between it and all current cluster centers c including c1 * ):
[0022]
[0023] In the formula, h is the number of current initial cluster centers, and d(·) is the Euclidean distance;
[0024] S13. Calculate the probability that each sample is selected as the next cluster center:
[0025]
[0026] Based on the calculated probability distribution, use the roulette wheel selection method to determine the next initial cluster center;
[0027] S14. Iteratively execute S12 and S13 until M initial cluster centers are selected Allocate the remaining non-cluster center samples to the corresponding clusters according to the nearest neighbor principle;
[0028] S15. Count the number of samples N contained in each cluster j , calculate the Euclidean distance between each initial cluster center c j and all samples y j,i in its cluster
[0029] Based on the Euclidean distance, further calculate the cluster radius of all clusters which is the Euclidean distance from the cluster center to the farthest sample in the cluster, and the calculation formula is as follows: r j = max(d(c j , y j,i ))), j = 1,..., M, i = 1,, N j .
[0030] For the unqualified clusters with a total number of P In step S3, based on the k-means++ clustering algorithm and the unqualified cluster determination rule, decompose each unqualified cluster into multiple qualified sub-clusters, and update M and the cluster structure information, including:
[0031] S31. Initialize u to 1 and Nsub to 2;
[0032] S32. Check whether u is greater than P. If it is greater, execute S33; if it is less, decompose the current cluster:
[0033] First, use the kmeans++ algorithm to re-cluster the samples within the cluster into Nsub sub-clusters. Then, use S2 to determine whether the sub-clusters are unqualified clusters. If all sub-clusters are qualified clusters, increment u by 1, initialize Nsub to 2, save the sub-cluster information, and then re-execute S32. Otherwise, increment Nsub by 1 and then re-execute S32.
[0034] S33. Update M and the cluster structure information by combining the sub-cluster information.
[0035] In step S4, the use of the triangle inequality checking strategy for K-nearest neighbor search of the sample to be queried includes:
[0036] S41. For the sample to be queried x, calculate the Euclidean distance between all cluster centers and x, and sort the distance calculation results in ascending order to obtain the sorted clusters where is the sorted cluster center;
[0037] If the number of samples in the first sorted cluster then search for the K-nearest neighbors of x by exhaustive search among all samples within , and at the same time set the index parameter m to 1, and the K-nearest neighbor set is where
[0038] Otherwise, construct a set that satisfies the following formula
[0039]
[0040] Search for the K-nearest neighbors of x by exhaustive search among all samples belonging to the cluster , and set d K to the maximum value of the Euclidean distances between all K-nearest neighbors and x;
[0041] S42. For other clusters outside the clusters that have been queried in S41, determine whether they may have potential nearest neighbors of x through the following formula:
[0042]
[0043] where is the cluster radius of the j-th sorted cluster ;
[0044] When the above formula does not hold, all samples in the cluster cannot become potential nearest neighbors of x and are directly excluded; if the above formula holds, it indicates that there may still be samples closer to the current K-nearest neighbors within the cluster , so for the cluster All samples in the test are subjected to triangular inequality judgment;
[0045] After all clusters have been identified by the triangle inequality check strategy, the final K nearest neighbors are output.
[0046] In step S42, the cluster All samples in the triangle inequality test are performed, including:
[0047] Cluster Heart Sample x to be queried, cluster Sample y in j,i Construct a triangle for which the triangle inequality holds absolutely true;
[0048] For sample y j,i , if equation (2) holds, then the triangle inequality and equation (2) are combined to derive equation (3), which means that y j,i The distance to x is farther than the current farthest neighbor, so y j,i It cannot be a neighbor of x;
[0049]
[0050] d(x,y j,i )>d K (3)
[0051] If equation (2) does not hold, then directly calculate d(x,y j,i ) to determine y j,i Is it better than the neighbor at this time? is closer than x. If it is closer, then Eliminate Then add y j,i And update the K nearest neighbor set and d K .
[0052] The Aurora optimization algorithm uses the inverse of the average distance calculation reduction ratio as the fitness function to optimize the number of clusters initially divided and the boundary adjustment factor.
[0053] The optimal values of the number of clusters M and the boundary adjustment factor λ obtained by the Aurora optimization algorithm include:
[0054] S51. For each set of hyperparameters M and λ to be optimized, the sample set is divided into A subsets, one of which is used as the sample set to be queried each time, and the remaining A-1 subsets are used as training sets. K nearest neighbor search is performed on each sample in the sample set to be queried through steps S1 to S4, and the corresponding RRDC is calculated:
[0055]
[0056] Among them, m is the number of times of directly calculating the Euclidean distance during the K-nearest neighbor search for all samples x to be queried, and Q is the number of samples to be queried;
[0057] Perform A calculations, and calculate the average distance calculation reduction ratio ARRDC based on the obtained A RRDCs:
[0058]
[0059] S52. Perform aurora optimization calculation, including:
[0060] Initialize FEs and it to 0, and set the maximum number of iterations MaxFEs;
[0061] Generate an initial high-energy particle swarm through the following formula, and each high-energy particle represents a group of candidate solutions of M and λ;
[0062] X(N H ,D H ) = LB + R × (UB - LB)
[0063] In the formula, N H is the number of high-energy particle swarms; D H is the scalable dimension of the solution space; LB and UB represent the lower and upper bounds of the solution space; R is a random number sequence in [0, 1];
[0064] For the initial population, calculate the fitness function f(X) = -ARRDC of each candidate solution, and use the candidate solution when f(X) is the smallest as the optimal solution X best ;
[0065] S53. For each initial high-energy particle, perform the following steps:
[0066] Calculate the velocity of the high-energy particle through the following formula:
[0067]
[0068] Among them, the iterative process is manifested as the movement process of the particle, that is, the FEs moment is the FEs-th iterative process; C is the integral constant, q is the charge carried by the particle, B is the Earth's magnetic field strength; m is the particle mass; α F is the damping factor;
[0069] Calculate the elliptical walk of each particle through the following formula:
[0070] Ao = Levy(D H ) × (X avg (j) - X(i, j)) + LB + r1 × (UB - LB) / 2
[0071] i = 1, 2,..., N H, j = 1, 2, ..., D H
[0072] Wherein, Levy(·) is Levy flight; X avg is the centroid position of the high-energy particle swarm; X(i, j) is the horizontal and vertical coordinates of the current position of the high-energy particle; r1 is a random sequence taking values in [0, 1];
[0073] The adaptive weight is calculated by the following two formulas:
[0074]
[0075] Update the position of the high-energy particle:
[0076] X new (i, j) = X(i, j) + r2 × (W1 × v(FEs) + W2 × Ao)
[0077] Wherein, r2 represents the environmental interference factor received by the particle, and its value range is [0, 1];
[0078] If r4 < K and r5 < 0.05, then update the position again through the following high-energy particle collision formula:
[0079] X new (i, j) = X(i, j) + sin(r3 × π) × (X(i, j) - X(a, j))
[0080] Wherein, X(a, j) represents any particle in the particle swarm; r3, r4, r5 are random values, and the value range is [0, 1];
[0081] Through the collision probability K c Control the increasingly frequent collisions between particles during the process of the algorithm:
[0082]
[0083] Obtain f(X new ) through the fitness function and increment FEs by 1;
[0084] When all high-energy particles have completed speed update and position adjustment, for X that satisfies f(X new ) < f(X), use the greedy selection algorithm for iterative update;
[0085] Reorder the iterated X according to f(X) and update the optimal solution X best After that, increment it by 1;
[0086] S54. If FEs ≥ MaxFEs, end the iteration and output X best And f(X best), X best is the optimized variable M best and λ best ; otherwise, execute S53.
[0087] In step S52, when generating the initial high-energy particle swarm, the upper and lower bounds of the solution space are set as follows: the boundary of λ is (1, 4], and the boundary of M is n is the number of samples in the sample set.
[0088] In step S51, the value of K in the K-nearest neighbor search process is set to 9.
[0089] The empirical setting value of the boundary adjustment factor λ is 3; the empirical setting value of the initially divided number of clusters M is n is the number of samples in the sample set.
[0090] The beneficial effects of the present invention are as follows:
[0091] The present invention significantly improves the search efficiency. In the offline stage, the present invention uses the kmeans++ algorithm to preliminarily divide the sample space, and proposes a reconstruction strategy based on the determination of unqualified clusters for the divided subspaces, performs a refined division of the sample space, divides the sample set into multiple subspaces, comprehensively and effectively processes some samples that affect the clustering quality, so that the samples in each subspace have higher similarity and a more compact spatial distribution, realizing a multi-level cluster structure division and optimization. Then, in the online stage, the sample space obtained by optimizing the cluster structure in the offline stage is searched through the triangle inequality strategy. Compared with traditional search methods, the search time in the high-dimensional space is significantly reduced, and the search time of the K-nearest neighbor is greatly reduced.
[0092] The present invention introduces the triangle inequality strategy in the online search process, effectively discriminates samples through the triangle inequality to screen potential nearest neighbor points, greatly reduces unnecessary direct distance calculations, and further reduces the time overhead.
[0093] The present invention innovatively uses the aurora optimization algorithm to adaptively optimize the number of initial clusters and the boundary adjustment factor, further improving the rationality of the space division, and helping to achieve efficient and accurate nearest neighbor search.
[0094] Other features and advantages of the present invention will be described in the subsequent specification or understood by implementing the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0095] Figure 1 is a schematic flowchart of the method of the embodiment of the present invention.
[0096] Figure 2Schematic diagram of the application principle of the triangle inequality in the embodiment of the present invention in the approximate nearest neighbor fast search algorithm.
[0097] Figure 3 Average RRDC change curve of different datasets in the embodiment of the present invention under different numbers of clusters at different K values.
[0098] Figure 4 TR change curve diagram at different K in the embodiment of the present invention. Detailed implementation manners
[0099] The following describes the detailed implementation manners of the present invention with reference to the accompanying drawings.
[0100] See Figure 1 , a K-nearest neighbor search method based on cluster structure optimization and triangle inequality strategy in the embodiment of the present application includes:
[0101] S1. Using the k-means++ clustering algorithm to preliminarily divide the sample set to be searched into M clusters with similar characteristics, and obtaining the cluster structure information of the sample set, which includes the cluster center position, cluster radius, number of samples, and cluster labels of each sample;
[0102] S2. Calculating the representative distance δ of each sample in each cluster j,i :
[0103]
[0104] where y j,i represents the i-th sample in the j-th cluster, Nj is the number of samples in y j,i , j = 1,..., M; d(y j,i , y j,q ) represents the Euclidean distance between y j,i and y j,q ;
[0105] Based on the representative distance, constructing a non-conforming cluster determination rule, and determining all clusters to obtain non-conforming clusters;
[0106] The determination rule includes:
[0107]
[0108] In the above formula, max(δ j,i ) is the maximum value of the representative distances of all samples in the cluster, λ represents a boundary adjustment factor, which is an empirically set value or the optimal value obtained through the aurora optimization algorithm;
[0109] If the samples in the cluster satisfy the above formula, it is determined as a conforming cluster, otherwise, it is determined as a non-conforming cluster;
[0110] Among them, the boundary adjustment factor λ and the initially divided number of clusters M are empirically set values or optimal values obtained through the aurora optimization algorithm;
[0111] S3. Based on the k-means++ clustering algorithm and the unqualified cluster determination rule, decompose each unqualified cluster into multiple qualified sub-clusters, update M and the cluster structure information, and obtain a sample set with updated structure;
[0112] S4. Based on the training sample set with updated structure, perform K-nearest neighbor search on the query sample using the triangle inequality checking strategy.
[0113] One of the core inventive points of this application is that after optimizing the cluster structure of the sample set first, then using the triangle inequality checking strategy to perform accurate K-nearest neighbor search in the sample space with optimized structure, which can significantly reduce the time consumption in the search process. It is applicable to the real-time nearest neighbor retrieval requirements in large-scale high-dimensional data scenarios. Among them, the cluster structure optimization is carried out in the offline stage, using the cluster structure optimization strategy to refine the division of the sample space, dividing the sample set into multiple sub-spaces, combining the determination of unqualified clusters, and effectively processing some samples that affect the clustering quality, so that the samples in each sub-space have higher similarity and more compact spatial distribution. The structure information of the sub-space will be used in the triangle inequality discrimination, and the time consumption in the discrimination process is reduced by storing information in advance.
[0114] Another core inventive point of this application is that the boundary adjustment factor λ and the initially divided number of clusters M can be the optimal values obtained through the aurora optimization algorithm. Among them, the number of clusters is a key parameter affecting the search performance. Too many clusters will increase the direct distance calculation between the query point and the cluster center, and too few clusters will result in low utilization efficiency of the cluster structure. In the process of dividing the search space, the initial number of clusters and the boundary adjustment factor are two factors affecting the number of clusters. Since the search space division process is carried out in the offline stage, the aurora optimization algorithm is selected to optimize the initial number of clusters and the boundary adjustment factor to further improve the search efficiency.
[0115] As a specific implementation, this embodiment is integrated into an industrial equipment status warning model based on the CMEW-EKNN algorithm. This model realizes fault prediction by analyzing the relationship between the real-time operation data of the equipment and the K-nearest neighbors in the sample library, so as to assist in evaluating the operation status of the equipment. And in the model architecture, the traditional K-nearest neighbor search algorithm and the method proposed in this embodiment are respectively used to compare the time efficiency ratio of the two, so as to illustrate the effect of this example. The sample set used is a 7-dimensional typical sample of a thermal power plant high-pressure heater containing 7,000 samples. The following describes the specific implementation.
[0116] S1. Use the k-means++ clustering algorithm to preliminarily divide the sample set to be searched into M clusters with similar characteristics, and obtain the cluster structure information of the sample set, including:
[0117] S11. For a sample set T composed of n = 7000 samples y ∈ R D select a sample randomly in T as the initial cluster center c1;
[0118] S12. For a non-cluster center sample y * , calculate the shortest distance D(y j ) between it and all current cluster centers c * including c1:
[0119]
[0120] where h is the current number of initial cluster centers and d(·) is the Euclidean distance;
[0121] S13. Calculate the probability that each sample is selected as the next cluster center:
[0122]
[0123] Based on the calculated probability distribution, use the roulette wheel selection method to determine the next initial cluster center; by simulating the rotation of the roulette wheel, make a random selection according to the probability values of each sample, which not only ensures the randomness of the selection but also ensures that instances with a greater distance have a higher probability of being selected, thus optimizing the distribution of the initial cluster centers;
[0124] S14. Iteratively execute S12 and S13 until M initial cluster centers are selected Allocate the remaining non-cluster center samples to the corresponding clusters using the nearest neighbor principle, that is, each non-cluster center sample is assigned to the cluster where the cluster center with the closest Euclidean distance to it is located;
[0125] where the initially divided number of clusters M is an empirically set value n is the number of samples in the sample set; or preferably the optimal value obtained through the aurora optimization algorithm;
[0126] S15. Count the number of samples N contained in each cluster j , calculate the Euclidean distance between each initial cluster center c j and all samples y j,i in its cluster i represents the index of the samples in the cluster;
[0127] Based on the Euclidean distance, further calculate the cluster radius of all clusters which is the Euclidean distance from the cluster center to the farthest sample in the cluster, and the calculation formula is as follows:
[0128] r j = max(d(c j , y j,i ), j = 1, ..., M, i = 1, ..., N j 。
[0129] Search spaces composed of different sample sets have different spatial characteristics. Using kmeans++ can better divide the search space into multiple sub-spaces called clusters with similar characteristics. The samples (i.e., instances) within these sub-spaces have similar characteristics and contain a large amount of cluster structure information. And the kmeans++ method is applicable to non-uniformly distributed search spaces. By obtaining cluster structure information in the offline stage before online search and using it in the formal search stage, the time consumed by the search can be effectively reduced.
[0130] S2. The spatial partitioning based on kmeans++ is relatively rough, which may result in sparse points within the cluster, and the sparse points may cause the cluster radius to be too large, leading to a decrease in the efficiency of the triangle inequality checking strategy. To address this, this step proposes a representative distance calculation formula for the sample set and constructs a non-conforming cluster determination rule based on the representative distance, which specifically includes the following operations:
[0131] S21. Calculate the representative distance δ of each sample within each cluster j,i :
[0132]
[0133] where y j,i represents the i-th sample in the j-th cluster, Nj is the number of samples in y j,i , j = 1, …, M; d(y j,i , y j,q ) represents the Euclidean distance between y j,i and y j,q ;
[0134] As can be seen from the above, the representative distance of the sample is defined as the sum of the Euclidean distances between the sample and other samples within the cluster to which it belongs, quantifying the tightness of the sample distribution within the cluster, characterizing the attribute characteristics of the sample in the cluster. The larger the representative distance, the lower the similarity between the sample and other samples in the cluster, and the greater the possibility of being a sparse point; conversely, the higher the similarity, the smaller the possibility of being a sparse point.
[0135] S22. Construct a non-conforming cluster determination rule based on the representative distance, determine all clusters, and obtain non-conforming clusters;
[0136] The determination rule is as follows: the overall characteristics of a cluster are characterized by the representative distance mean of the samples within the cluster. If the representative distance mean of all samples within a cluster multiplied by the boundary adjustment factor λ (λ > 1, and its value is related to the characteristics of the sample set) is greater than the maximum value of the representative distance, that is, when the following formula holds, the cluster is determined to be a qualified cluster; conversely, if the following formula does not hold, the cluster is determined to be an unqualified cluster;
[0137]
[0138] In the above formula, max(δ j,i ) is the maximum value of the representative distance of all samples within the cluster;
[0139] Among them, the boundary adjustment factor λ can be an empirically set value of 3, or preferably the optimal value obtained through the aurora optimization algorithm;
[0140] Represent all unqualified clusters as and the number thereof is P;
[0141] S3. For the P unqualified clusters Based on the k-means++ clustering algorithm and the unqualified cluster determination rule, each unqualified cluster is decomposed into multiple qualified sub-clusters, and M and the cluster structure information are updated, including:
[0142] S31. Initialize u to 1 and Nsub to 2;
[0143] S32. Check whether u is greater than P. If it is greater, execute S33; if it is less, decompose the current cluster:
[0144] First, re-cluster the samples within the cluster into Nsub sub-clusters through the kmeans++ algorithm. Then, perform unqualified cluster discrimination on the sub-clusters through S2; if all sub-clusters are qualified clusters, let u be incremented by 1 and Nsub be initialized to 2, and after saving the sub-cluster information, re-execute S32; otherwise, let Nsub be incremented by 1 and then re-execute S32;
[0145] S33. Update M in combination with the sub-cluster information, and execute S15 to update the cluster structure information.
[0146] Among them, an unqualified cluster represents that there are sparse points within the cluster, and the sparse points will cause the cluster radius to be too large, thereby reducing the algorithm performance. Step S3 is based on the unqualified cluster determination rule to perform iterative decomposition and determination on the unqualified clusters through kmeans++, decompose them into multiple qualified small clusters, further refine the search space, improve the utilization efficiency of the space structure information, realize the reconstruction of unqualified clusters, thereby optimizing the cluster structure and improving the clustering quality.
[0147] S4. Adopt the triangular inequality checking strategy to perform K-nearest neighbor search on the query sample in the sample set updated in S3, including:
[0148] S41. For the query sample x, calculate the Euclidean distances between all cluster centers and x, and sort the distance calculation results in ascending order to obtain the sorted clusters where, is the sorted cluster center;
[0149] If the number of samples in the first sorted cluster is then search for the K-nearest neighbors of x by exhaustive search among all samples within , and at the same time set the index parameter m to 1, and the K-nearest neighbor set is where
[0150] Otherwise, construct a set that satisfies the following formula
[0151]
[0152] Search for the K-nearest neighbors of x by exhaustive search among all samples belonging to the cluster , and set d K to the maximum value of the Euclidean distances between all K-nearest neighbors and x;
[0153] That is, after step S41, for the cluster index m, it satisfies: the sum of the number of instances in the first m clusters is greater than or equal to K, and the sum of the number of instances in the first m - 1 clusters is less than K. Calculate the Euclidean distances between the query sample and all instances in the first m clusters, sort them, select the nearest K instances as the current K-nearest neighbors, and use the distance from the current farthest neighbor to the query sample as the neighbor radius d K ;
[0154] S42. For other clusters outside the clusters that have been queried in S41, that is, for clusters in the ordered cluster set of S41 with an index greater than m, judge whether they may have potential neighbors of x through the following formula:
[0155]
[0156] where, is the cluster radius of the j-th sorted cluster ;
[0157] When the above formula does not hold, all samples in the cluster cannot become potential neighbors of x and are directly excluded; if the above formula holds, it indicates that there may still be samples closer to the current K-nearest neighbors within the cluster , so for the cluster Perform the triangle inequality discrimination on all samples within. That is, if a circle with the current cluster center as the center and the cluster radius as the radius intersects with a circle with the sample to be queried as the center and d K as the radius, then there are potential near neighbors in the current cluster; otherwise, there are no potential near neighbors. For the instances in the cluster with potential near neighbors, further screening is performed using the triangle inequality checking strategy;
[0158] After all clusters have been discriminated by the triangle inequality checking strategy, the final K nearest neighbors are output.
[0159] Among them, the triangle inequality discrimination on all samples within the cluster includes:
[0160] Taking the cluster center the sample x to be queried, and the sample y within the cluster to construct a triangle as j,i shown. For this triangle, the triangle inequality (1) definitely holds: Figure 2
[0161]
[0162] For the sample y in , if equation (2) holds, then by combining the triangle inequality (1) and equation (2), equation (3) is derived, which means that the distance from y j,i to x is farther than the current farthest neighbor, so y j,i cannot be a near neighbor of x; j,i
[0163]
[0164] d(x, y j,i ) > d K (3)
[0165] If equation (2) does not hold, then directly calculate d(x, y j,i ) to determine whether y j,i is closer to x than the current near neighbor . If it is closer, then remove from and add y and update the current K nearest neighbor set and d j,i . K As a preferred implementation manner, the method of this embodiment further includes: S5. Based on the aurora optimization algorithm, using the opposite of the reduction ratio calculated by the average distance as the fitness function, optimize the initially divided number of clusters and the boundary adjustment factor, specifically including the following steps:
[0166]
[0167] S51. Obtain the average distance calculation reduction ratio ARRDC through ten-fold cross-validation:
[0168] For each set of hyperparameters M and λ to be optimized, divide the sample set into A subsets. Each time, use one of the subsets as the sample set to be queried, and the remaining A - 1 subsets as the training set. Perform K-nearest neighbor search on each sample in the sample set to be queried through steps S1 to S4, and calculate the corresponding RRDC:
[0169]
[0170] Among them, m is the number of times the Euclidean distance is directly calculated for all samples x to be queried during the K-nearest neighbor search process, and Q is the number of samples to be queried;
[0171] Perform A calculations, and calculate the average distance calculation reduction ratio ARRDC based on the obtained A RRDC values:
[0172]
[0173] A is preferably 10; RRDC i represents the RRDC calculated during the search process when the i-th subset is used as the sample set to be queried;
[0174] Figure 3 Shows the RRDC change curves for different numbers of clusters at different K values under ten-fold cross-validation for four different datasets (Abalone dataset, Orohd dataset, Lr dataset, Rice dataset). It can be seen from the figure that the change trends of the curves are basically the same for different K values, indicating that K has a relatively small impact on the optimal variables. Therefore, in this embodiment, K is set to 9;
[0175] S52. Perform aurora optimization calculation, including:
[0176] Initialize FEs and it to 0, and set the maximum number of iterations MaxFEs;
[0177] Generate an initial high-energy particle swarm through the following formula, where each high-energy particle represents a candidate solution for a set of M and λ;
[0178] X(N H , D H ) = LB + R × (UB - LB)
[0179] In the formula, N H is the number of high-energy particles in the swarm, generally set to 30 - 100; D H is the scalable dimension of the solution space, related to the number of optimization variables. Here, D H is set to 2; LB and UB represent the lower and upper bounds of the solution space. The boundary of λ is preferably (1, 4], and the boundary of M is preferably n is the number of samples in the sample set; R is a random number sequence in [0, 1];
[0180] For the initial population, calculate the fitness function of each candidate solution f(X) = -ARRDC, and take the candidate solution when f(X) is the smallest as the optimal solution X best ;
[0181] S53. For each initial high-energy particle, perform the following steps:
[0182] Calculate the velocity of the high-energy particle through the following formula:
[0183]
[0184] where the iterative process is manifested as the movement process of the particle, that is, the FEs moment is the FEs-th iterative process; C is the integral constant, q is the charge carried by the particle, and B is the Earth's magnetic field strength, all of which are taken as 1; m is the particle mass, taken as 100; α F is the damping factor, and its value range is [1, 1.5].
[0185] Calculate the elliptical walk of each particle through the following formula:
[0186] Ao = Levy(D H )×(X avg (j)-X(i,j))+LB+r1×(UB-LB) / 2
[0187] i = 1, 2,..., N H , j = 1, 2,..., D H
[0188] In the formula, Levy(·) is the Levy flight, which is a random non-Gaussian walk; X avg is the centroid position of the high-energy particle swarm; X(i, j) is the current position of the high-energy particle, and it can be understood that i and j represent the horizontal and vertical coordinates respectively; r1 is a random sequence taking values in [0, 1];
[0189] Calculate the adaptive weight through the following two formulas:
[0190]
[0191] Perform the following three steps for each high-energy particle:
[0192] (1) Based on the above velocity formula, elliptical walk formula set and adaptive weight formula calculation, update the position of each high-energy particle:
[0193] X new (i,j) = X(i,j)+r2×(W1×v(FEs)+W2×Ao)
[0194] Wherein, r2 represents the environmental interference factor suffered by the particle, and its value range is [0, 1];
[0195] (2) If r4 < K and r5 < 0.05, then update the position again through the following high-energy particle collision formula:
[0196] X new (i,j) = X(i,j) + sin(r3 × π) × (X(i,j) - X(a,j))
[0197] Wherein, X(a,j) represents any particle in the particle swarm; r3, r4, and r5 are random values, and their value range is [0, 1];
[0198] Through the collision probability K c Control the increasingly frequent collisions between particles during the process of the algorithm:
[0199]
[0200] (3) Obtain f(X new ) through the fitness function and increment FEs by 1.
[0201] When all high-energy particles have completed velocity update and position adjustment, for X that satisfies f(X new ) < f(X), use the greedy selection algorithm for iterative update: that is, for each high-energy particle, if the fitness value f(X new ) of its new position X new ) is less than the fitness value f(X) of the current position X, then accept the new position as the current solution; otherwise, retain the original position X. Reorder X after iteration according to f(X) and update the optimal solution X best After that, increment it by 1;
[0202] S54. If FEs ≥ MaxFEs at this time, end the iteration and output X best and f(X best ), X best is the optimized variable M best and λ best ; otherwise, execute S53.
[0203] Since steps S1 to S3 and S5 of this embodiment all belong to the stage of offline extracting cluster structure information and optimizing parameters, and all steps are completed before the online search stage of S4, the time consumption during online search is effectively saved.
[0204] In the real-time warning of the CMEW-EKNN state warning model, the time efficiency ratio, i.e., the TR index, is used to illustrate the performance advantage of the method of this embodiment over the traditional K-nearest neighbor search method. The TR index is defined as the ratio of the time spent by CMEW-EKNN with the traditional K-nearest neighbor search as the architecture in state warning to the time consumed with the architecture of the present invention. The larger this ratio is, the less time-consuming the method of this embodiment is in the same search task. As Figure 4 shown, the test data for 1000 groups of real-time device state samples shows that when the nearest neighbor parameter K takes 9, 55, and 101, the TR values reach 5, 4.4, and 3.1 respectively, verifying that the method of this embodiment has a significant acceleration effect in different-scale nearest neighbor search scenarios, and shows the best time efficiency improvement when the K value is small (K = 9).
[0205] In summary, the present invention is applicable to fast search under the conditions of large-scale high-dimensional data sets, obtains historical operation data samples that conform to the current operation state of the device, and can meet the device fault diagnosis requirements in complex industrial scenarios.
[0206] Those of ordinary skill in the art can understand that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A K-nearest neighbor search method based on cluster structure optimization and triangle inequality strategy, characterized in that: include: S1. Using the k-means++ clustering algorithm, the sample set to be searched is preliminarily divided into M clusters with similar characteristics, and the cluster structure information of the sample set is obtained, which includes the cluster center position, cluster radius, number of samples and cluster label of each sample; S2. Calculate the representative distance δ of each sample in each cluster j,i : Among them, y j,i represents the i-th sample in the j-th cluster, N j is the number of samples in the jth cluster, j = 1, ..., M; d(y j,i ,y j,q ) represents y j,i With y j,q The Euclidean distance of Constructing a non-conforming cluster determination rule based on the representative distance, determining all clusters, and obtaining non-conforming clusters; The determination rules include: In the above formula, max(δ j,i ) is the maximum value of the representative distance of all samples in the cluster, λ is the boundary adjustment factor, which is the empirical setting value or the optimal value obtained by the Aurora optimization algorithm; if the samples in the cluster meet the above formula, they are determined to be qualified clusters, otherwise, they are determined to be unqualified clusters; Among them, the boundary adjustment factor λ and the number of clusters M divided initially are empirically set values or the optimal values obtained by the Aurora optimization algorithm; S3, based on the k-means++ clustering algorithm and the unqualified cluster determination rule, decompose each unqualified cluster into a plurality of qualified sub-clusters, update M and the cluster structure information, and obtain a sample set with updated structure; S4. Based on the training sample set after the structural update, a triangle inequality check strategy is used to perform a K-nearest neighbor search on the query sample.
2. The method according to claim 1, characterized in that In step S1, the k-means++ clustering algorithm is used to preliminarily divide the sample set to be searched into M clusters with similar characteristics to obtain cluster structure information of the sample set, including: S11. For n samples y∈R D The sample set T is constructed, and a sample is randomly selected from T as the initial cluster center c1; S12, for non-cluster center sample y * , calculate its relationship with all current cluster centers c including c1 j The shortest distance D(y * ): In the formula, h is the number of current initial cluster centers, d(·) is the Euclidean distance; S13. Calculate the probability of each sample being selected as the next cluster center: Based on the calculated probability distribution, the roulette wheel selection method is used to determine the next initial cluster center; S14, iteratively execute S12 and S13 until M initial cluster centers are selected The remaining non-cluster center samples are assigned to the corresponding clusters using the nearest assignment principle; S15. Count the number of samples N contained in each cluster j , calculate the initial cluster center c j All samples y in its cluster j,i The Euclidean distance between Based on the Euclidean distance, the cluster radius of all clusters is further calculated It is the Euclidean distance from the cluster center to the farthest sample in the cluster, and the calculation formula is as follows: r j =max(d(c j ,y j,i )),j=1,...,M,i=1,...,N j 。 3. The method according to claim 1, characterized in that For a total number of P unqualified clusters In step S3, based on the k-means++ clustering algorithm and the unqualified cluster determination rule, each unqualified cluster is decomposed into a plurality of qualified sub-clusters, and M and the cluster structure information are updated, including: S31, initialize u to 1, initialize Nsub to 2; S32, check whether u is greater than P, if so, execute S33; if less, decompose the current cluster: First, the samples in the cluster are re-clustered into Nsub sub-clusters through the kmeans++ algorithm, and then the sub-clusters are identified as unqualified clusters through S2; if all sub-clusters are qualified clusters, u is incremented by 1 and Nsub is initialized to 2, and the sub-cluster information is saved and S32 is re-executed; otherwise, Nsub is incremented by 1 and S32 is re-executed; S33. Update M and the cluster structure information in combination with the sub-cluster information.
4. The method according to claim 1, characterized in that: In step S4, the use of the triangle inequality check strategy to perform K nearest neighbor search on the query sample includes: S41. For the query sample x, calculate the Euclidean distance between all cluster centers and x, and sort the distance calculation results in ascending order to obtain the sorted clusters. in, is the cluster center after sorting; If the first cluster after sorting The number of samples in In The K nearest neighbors of x are searched by exhaustive method in all samples in, and the index parameter m is set to 1. The set of K nearest neighbors is in Otherwise, construct a set that satisfies the following formula In the By exhaustively searching for the K nearest neighbors of x in all samples of the cluster, K Set it to the maximum value of the Euclidean distance between all K nearest neighbors and x; S42. For other clusters other than the cluster that has been queried in S41, determine whether they may have potential neighbors of x by the following formula: in, is the jth cluster after sorting The cluster radius of When the above formula does not hold, the cluster All samples in cannot become potential neighbors of x and are directly excluded; if the above formula holds, it means that cluster There may still be samples closer than the current K nearest neighbors in the cluster. All samples in the test are subjected to triangular inequality judgment; After all clusters have been identified by the triangle inequality check strategy, the final K nearest neighbors are output.
5. The method according to claim 4, characterized in that In step S42, the cluster All samples in the triangle inequality test are performed, including: Cluster Heart Sample x to be queried, cluster Sample y in j,i Construct a triangle for which the triangle inequality holds absolutely true; For sample y j,i , if equation (2) holds, then the triangle inequality and equation (2) are combined to derive equation (3), which means that y j,i The distance to x is farther than the current farthest neighbor, so y j,i It cannot be a neighbor of x; If equation (2) does not hold, then directly calculate d(x,y j,i ) to determine y j,i Is it better than the neighbor at this time? is closer than x. If it is closer, then Eliminate Then add y j,i And update the K nearest neighbor set and d K .
6. The method according to claim 1, characterized in that The Aurora optimization algorithm uses the inverse of the average distance calculation reduction ratio as the fitness function to optimize the number of clusters initially divided and the boundary adjustment factor.
7. The method according to claim 1 or 6, characterized in that: The optimal values of the number of clusters M and the boundary adjustment factor λ obtained by the Aurora optimization algorithm include: S51. For each set of hyperparameters M and λ to be optimized, the sample set is divided into A subsets, one of which is used as the sample set to be queried each time, and the remaining A-1 subsets are used as training sets. K nearest neighbor search is performed on each sample in the sample set to be queried through steps S1 to S4, and the corresponding RRDC is calculated: Where m is the number of times all query samples x are directly calculated for Euclidean distance in the K nearest neighbor search process, and Q is the number of query samples; Perform A calculations, and calculate the average distance calculation reduction ratio ARRDC based on the A RRDCs obtained: S52, performing aurora optimization calculation, including: Initialize FEs and it to 0, and set the maximum number of iterations MaxFEs; The initial high-energy particle group is generated by the following formula, where each high-energy particle represents a set of candidate solutions for M and λ; X(N H ,D H )=LB+R×(UB-LB) Where N H is the number of high-energy particle groups; D H is the scalable dimension of the solution space; LB and UB represent the lower and upper bounds of the solution space; R is a random number sequence of [0,1]; For the initial population, calculate the fitness function f(X) = -ARRDC for each candidate solution, and take the candidate solution with the minimum f(X) as the optimal solution X best ; S53. For each initial high-energy particle, perform the following steps: The speed of high-energy particles is calculated by the following formula: The iteration process is represented by the movement process of the particle, that is, the FEs moment is the FEsth iteration process; C is the integral constant, q is the charge carried by the particle, B is the earth's magnetic field strength; m is the particle mass; α F is the damping factor; The elliptical walk of each particle is calculated by the following formula: Ao=Levy(D H )×(X avg (j)-X(i,j))+LB+r1×(UB-LB) / 2 i=1,2,...,N H ,j=1,2,...,D H Where, Levy(·) is Levy flight; X avg is the center of mass of the high-energy particle group; X(i,j) is the horizontal and vertical coordinates of the current position of the high-energy particle; r1 is a random sequence with a value between [0,1]; The adaptive weight is calculated by the following two formulas: Update the position of the energetic particles: X new (i,j)=X(i,j)+r2×(W1×v(FEs)+W2×Ao) In the formula, r2 represents the environmental interference factor to which the particle is subjected, and its value range is [0,1]; If r4<K and r5<0.05, the position is updated again by the following high-energy particle collision formula: X new (i,j)=X(i,j)+sin(r3×π)×(X(i,j)-X(a,j)) In the formula, X(a,j) represents any particle in the particle swarm; r3, r4, r5 are random values ranging from [0,1]; By the collision probability K c The control algorithm tends to cause frequent collisions between particles: By using the fitness function to obtain f(X new ) and increment FEs by 1; After all high-energy particles have completed velocity update and position adjustment, for X that satisfies f(X new ) < f(X), iterative update is performed using the greedy selection algorithm; Reorder the iterated X according to f(X) and update the optimal solution X best Then, let it add 1; S54, if FEs ≥ MaxFEs, end the iteration and output X best and f(X best ), X best This is the optimized variable M best With λ best ; Otherwise, execute S53.
8. The method according to claim 7, characterized in that In step S52, when generating the initial high-energy particle swarm, the upper and lower bounds of the solution space are set as follows: the boundary of λ is (1,5], the boundary of M is n is the number of samples in the sample set.
9. The method according to claim 7, characterized in that: In step S51, the K value in the K nearest neighbor search process is set to 9.
10. The method according to claim 1, characterized in that The empirical setting value of the boundary adjustment factor λ is 3; the empirical setting value of the number of clusters M divided initially is n is the number of samples in the sample set.
Citation Information
Cited By
High-dimensional medical feature selection method, system and device and storage medium
CN120527024A