Fast large-scale image data preprocessing method and application based on bipartite graph cut
Through the method based on the two-part graph cutting, the in-class similarity between the anchor point layer and the sample layer is directly optimized, which solves the problem of information loss in the prior art, and realizes efficient preprocessing and precise clustering of large-scale image data.
Patent Information
- Application Number
- CN202310419562.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-19
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-04-19
AI Technical Summary
The existing two-part graph learning clustering method causes the loss of key information due to the relaxation and discretization process, resulting in poor image preprocessing accuracy and inability to accurately characterize the original data cluster structure.
Using a two-part graph cutting method, by maximizing the in-class similarity between the anchor point layer and the sample layer, the discrete indication matrix is directly used to optimize the two-part graph cutting problem, avoiding the singular value decomposition relaxation process and discrete post-processing process, and reducing clustering time by using discrete coordinate rise optimization algorithm.
The accuracy and efficiency of image data preprocessing have been significantly improved, clustering accuracy and normalized mutual information have been improved by 2.64 and 3.11 percentage points respectively, and the clustering time has been reduced from 227.71 seconds to 23.14 seconds.
Smart Images

Figure CN116524221B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of image recognition and classification and pattern recognition, and relates to a fast large-scale image data preprocessing method and application based on bipartite graph cut Background Art
[0002] The rapid development of Internet and communication technologies in recent years has put forward higher requirements for the processing of data volume. Since the cumbersome label acquisition process is avoided, clustering technology is widely used in data processing scenarios such as virus spread prediction, incomplete multi-view filling, and time series monitoring in an unsupervised manner. Among them, in the field of image retrieval, researchers often first perform preprocessing operations on existing image data through clustering technology. When new samples outside the existing image data need to be identified, the nearest clustering center to the new sample is found through image feature information and attributed to the cluster, and then the most similar sample to the new sample is further found within the cluster to complete the accurate image retrieval task. The above image retrieval steps avoid the operation of comparing new samples with existing samples one by one, and significantly improve the retrieval efficiency of large-scale data
[0003] As the most classic graph learning algorithm, spectral clustering is usually used in the preprocessing step of big data engineering image retrieval due to its simplicity and ease of operation. It relaxes the binary indicator matrix into a continuous orthogonal matrix because it cannot handle the NP-hard problem caused by discrete optimization. The specific steps are as follows: 1) Construct a full-sample neighbor graph through a kernel function and obtain the Laplacian matrix; 2) Perform eigenvalue decomposition on the Laplacian matrix to obtain a low-dimensional embedding indicator matrix; 3) Perform discrete operations such as K-means and spectral rotation on the low-dimensional indicator matrix to obtain full-sample clustering labels. However, the time complexity of the graph construction and eigenvalue decomposition operations is proportional to the square and cube of the number of samples respectively, which severely restricts the application of spectral clustering in large-scale data preprocessing tasks
[0004] Therefore, many scholars have used anchor point downsampling learning strategies to improve the inefficiency of traditional spectral clustering. For example, Luo Xinglong et al. (Fast Spectral Clustering with Binary K-means Anchor Point Extraction, Computer Engineering and Applications, 2023, 1-10.) proposed a fast spectral clustering method based on anchor graph learning, which downsamples the original data through the binary k-means method and generates multi-layer anchor points, and constructs a hierarchical bipartite graph between the original data sample and the last layer of anchor points to reduce the scale of the composition. Subsequently, the method converts the traditional graph eigenvalue decomposition into a small-scale bipartite graph singular value decomposition to improve clustering efficiency. Nevertheless, as a variant of spectral clustering, the "relaxation-discretization" two-step solution strategy of this model will still affect the clustering accuracy due to the loss of key information, and its actual effect is worse than the model in the present invention. For example, when processing large-scale image data MNIST, the clustering accuracy and normalized mutual information of the fast spectral clustering method of binary k-means anchor extraction proposed by Luo Xinglong et al. were only 66.42% and 67.79%, and the algorithm single running time was 227.71 seconds. The method proposed in the present invention can complete the preprocessing task in 23.14 seconds and achieve 69.10% and 70.90% accuracy and normalized mutual information, which are 2.64 percentage points and 3.11 percentage points higher, respectively, and the clustering efficiency has been significantly improved. Among them, clustering accuracy and normalized mutual information are commonly used indicators for evaluating clustering performance. The larger the value, the better the clustering performance. Summary of the invention
[0005] Technical issues to be solved
[0006] In order to avoid the shortcomings of the prior art, the present invention proposes a fast large-scale image data preprocessing method and application based on bipartite graph cuts. The existing bipartite graph learning clustering method still predicts sample labels through a two-step strategy of first relaxation and then discretization. Due to the loss of key information, it is usually unable to accurately characterize the original data cluster structure, which leads to the generally poor accuracy of image preprocessing.
[0007] The present invention proposes an efficient discrete clustering algorithm for large-scale image data for data preprocessing before large-scale image retrieval, which is called a fast large-scale image data preprocessing method based on bipartite graph cut. This method uses a fixed sparse bipartite graph matrix to maximize the intra-class similarity between the anchor layer and the sample layer, and directly uses discrete indicator matrix optimization to solve the bipartite graph cut problem. This method not only avoids the singular value decomposition relaxation process and discrete post-processing process, reduces information loss, but also greatly reduces the clustering time through the discrete coordinate ascent optimization algorithm, thereby achieving high efficiency of large-scale data preprocessing.
[0008] Technical Solution
[0009] A fast large-scale image data preprocessing method based on bipartite graph cut, characterized by the following steps:
[0010] Step 1: Stretch and integrate n images of size a×b into an image data matrix where d = a×b is the number of pixels of a single image data;
[0011] Step 2: Perform downsampling on the image data matrix X, that is, use the hierarchical bipartite K-means algorithm to select m = 2 a anchor points, and obtain the anchor point data matrix where: a is the number of layers of the binary tree of the hierarchical bipartite K-means algorithm;
[0012] Construct the following bipartite graph problem:
[0013]
[0014] where, b i is the i-th row vector of the bipartite graph matrix B, representing the bipartite graph similarity between the i-th image vector x i and all other image vectors, 1 m is a column vector with all m elements being 1; b ij is the element in the i-th row and j-th column of B, representing the bipartite graph similarity between the i-th image vector x i and the j-th anchor point vector z j ; γ is the regularization parameter of the sparse regularization term;
[0015] Solve the above bipartite graph problem through the following formula, and calculate and obtain the bipartite graph matrix with the number of nearest neighbors being r
[0016]
[0017] where, r is the number of anchor point nearest neighbors of each sample;
[0018] Step 3: Respectively construct the discrete image sample label matrix and the image anchor point label matrix and define f jk as the element in the j-th row and k-th column of the matrix F, g jk as the element in the j-th row and k-th column of the matrix G, c is the number of categories included in the image data set; based on the matrices F, G, and B, construct the following within-class similarity maximization discrete model for the anchor point layer and the sample layer:
[0019]
[0020] At the same time, perform orthogonal normalization on F and G to obtain the discrete normalization model as follows:
[0021]
[0022] Then, initialize F and G randomly according to the above construction rules respectively;
[0023] Step 4: Optimize the discrete normalization model constructed in Step 3 by alternately iteratively updating F and G:
[0024] 1) Fix F and update G:
[0025] Calculate and calculate the incremental matrix T row by row according to the following formula:
[0026]
[0027] where, v k and g k are the k-th column vectors of V and G respectively, v jk and g jk are the j-th elements of the vectors v k and g k respectively, t jk is the element in the j-th row and k-th column of the incremental matrix T, representing the incremental value of the objective function corresponding to the change of the k-th class label index of the j-th image anchor point;
[0028] Then, based on the incremental matrix T, update the indicator matrix G according to the following formula:
[0029]
[0030] where, <·> is a logical indicator, with a value of 1 if the logic is true, otherwise 0; is the updated label element of the j-th image anchor point with respect to the k-th class;
[0031] 2) Fix G and update F:
[0032] Calculate and calculate the incremental matrix T * row by row according to the following formula:
[0033]
[0034] where, u k and f k are the k-th column vectors of U and F respectively, u ik and f ik are the i-th elements of the vectors u k and f k respectively, is the element in the i-th row and k-th column of the incremental matrix T * , representing the incremental value of the objective function corresponding to the change of the k-th class label index of the i-th image sample point;
[0035] Then, based on the incremental matrix T * update the indication matrix F according to the following formula:
[0036]
[0037] wherein, is the updated label element of the i-th image sample point with respect to the k-th class;
[0038] 3) Calculate the objective function value of the discrete normalization model, and judge whether it converges by the difference from the previous objective function value:
[0039] If it does not converge, return to ①; if it converges, go to step 5;
[0040] Step 5: Output F and G to obtain the predicted labels of n original image samples and m image anchors respectively. At this time, the image data preprocessing process ends.
[0041] The regularization parameter γ can be adaptively determined during the solution process of the bipartite graph problem.
[0042] The construction rules of the matrices F and G are as follows: if the i-th image sample belongs to the k-th class, then f ik = 1, otherwise f ik = 0, to represent whether x i belongs to the k-th class. If the j-th image anchor belongs to the k-th class, then g jk = 1, otherwise g ik = 0, to represent whether z j belongs to the k-th class.
[0043] An application of the fast large-scale image data preprocessing method based on bipartite graph cut, characterized in that: after the preprocessing process is completed, fast retrieval of large-scale image data can be realized by using full-sample clustering center comparison or image anchor comparison.
[0044] The fast retrieval strategy for the large-scale image data is as follows: 1) Calculate the clustering center of each cluster according to the labels of all original image data samples. When introducing new image data, first compare it with c clustering centers, assign it to the cluster of the nearest clustering center, and then further find the most similar original image sample in the corresponding cluster; 2) First calculate the labels of all anchor images by using G, and then directly compare the introduced new image with all anchors and assign it to the corresponding cluster of the nearest anchor to achieve retrieval.
[0045] Beneficial effects
[0046] A fast large-scale image data preprocessing method and application based on bipartite graph cut proposed by the present invention elongates and integrates n images of a×b scale into an image data matrix Perform downsampling on the image data matrix X to obtain a bipartite graph matrix Orthogonal normalization is performed on F and G to obtain a discrete normalization model. The discrete normalization model constructed in step 3 is optimized by alternately iteratively updating F and G; output F and G to obtain the predicted labels of n original image samples and m image anchor points respectively, and the image data preprocessing process is completed at this time. After the preprocessing process is completed, the full sample cluster center comparison or image anchor point comparison can be used to achieve large-scale out-of-sample new image data retrieval. With the help of a fixed sparse bipartite graph matrix, this method directly uses a discrete indicator matrix to optimize and solve the bipartite graph cut problem by maximizing the intra-class similarity between the anchor point layer and the sample layer. This method not only avoids the singular value decomposition relaxation process and the discrete post-processing process, reduces information loss, but also greatly reduces the clustering time through the discrete coordinate ascent optimization algorithm, thereby achieving high efficiency in large-scale data preprocessing.
[0047] (1) The model proposed in this paper aims to maximize the intra-class similarity between the anchor point layer and the sample layer, directly deal with the variation of the original graph cut problem of spectral clustering, namely the bipartite graph cut problem, and improve the accuracy of image data preprocessing through the interaction between anchor points and sample labels.
[0048] (2) The model proposed in the present invention does not contain any other hyperparameters except integer parameters, namely, the number of anchor points and the number of bipartite graph neighbors. Since the number of anchor points and the number of bipartite graph neighbors can usually be given based on experience in bipartite graph learning, the present invention simplifies the clustering model and greatly reduces the burden of parameter adjustment in practical applications.
[0049] (3) The proposed method directly performs iterative coordinate ascent optimization on discrete labels, thus avoiding the loss of key information caused by data post-processing such as K-means or spectral rotation. In addition, its time complexity is linearly related to the number of image data samples. Compared with traditional spectral clustering, it can not only significantly improve the efficiency of image data preprocessing, but also speed up data retrieval with the help of discrete anchor point labels.
[0050] (4) The method of the present invention is used for retrieval. When processing the large-scale image data MNIST, the clustering accuracy and normalized mutual information of the fast spectral clustering method of binary k-means anchor extraction proposed by Luo Xinglong et al. are only 66.42% and 67.79%, and the single algorithm running time is 227.71 seconds. Compared with the result, the method proposed in the present invention can complete the preprocessing task in 23.14 seconds and achieve 69.10% and 70.90% accuracy and normalized mutual information, which are 2.64 percentage points and 3.11 percentage points higher, respectively, and the clustering efficiency is significantly improved. Among them, clustering accuracy and normalized mutual information are commonly used indicators for evaluating clustering performance. The larger the value, the better the clustering performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 : Flowchart of image data preprocessing for the method proposed by the present invention
[0052] Figure 2 : Variation of the objective function value and clustering accuracy with the number of iterations in a single experiment of the MNIST image dataset Detailed implementation manners
[0053] The present invention will be further described in conjunction with embodiments and the accompanying drawings as follows:
[0054] The present invention proposes a fast large-scale image data preprocessing method based on bipartite graph cut. Taking the handwritten digit image dataset MNIST as an example, the specific implementation manners of performing data preprocessing for the proposed method are described, but the technical content of the present invention is not limited to the described scope. The dataset MNIST contains a total of 70,000 handwritten digit images of 10 categories with a pixel scale of 28×28, and the specific implementation steps are as follows:
[0055] Step 1: Stretch the 70,000×24×24 scale image dataset into a data matrix where 70,000 is the number of images, and 784 = 24×24 is the total number of pixels of a single image. This step aims to merge the original image data into a data matrix for subsequent matrix operations.
[0056] Step 2: Downsample the image data matrix X obtained in the previous step. First, use the hierarchical bipartite K-means algorithm to obtain 2 a = m (m << 70,000) anchor points that can roughly cover the original image data structure and generate an anchor point data matrix where a is the number of layers of the binary tree of the hierarchical bipartite K-means algorithm; then adaptively construct a sparse bipartite graph matrix by solving the following problem:
[0057]
[0058] where, b i is the bipartite graph vector corresponding to the i-th image vector in the bipartite graph matrix , b ij is the bipartite graph similarity between the i-th image vector x i and the j-th anchor point vector z j , γ is the regularization parameter of the sparse regularization term, 1 m is a vector with all elements of m dimensions being 1. The above sparse bipartite graph matrix problem has the following closed-form solution:
[0059]
[0060] where, r is the number of anchor point neighbors of each sample; the regularization parameter γ can be adaptively determined during the solution process.
[0061] Step 3: Based on the sparse bipartite graph matrix B, inspired by the original graph cut problem of spectral clustering (maximizing the within-class similarity of the entire sample matrix), a discrete model for maximizing the within-class similarity between the anchor layer and the sample layer is constructed. First, the formulation of this discrete model is as follows:
[0062]
[0063] where c is the number of true classes contained in the original image matrix, which is 10 for the MNIST dataset. If the i-th image sample belongs to the k-th class, then f ik = 1; otherwise, f ik = 0. Similarly, if the j-th image anchor belongs to the k-th class, then g jk = 1; otherwise, g ik = 0. According to the above discrete model, when and only when f ik = g jk = 1, the bipartite graph similarity b ij will be regarded as an effective element for maximizing the within-class similarity optimization. Therefore, this model minimizes the similarity loss between the anchor layer and the sample layer through bipartite graph similarity pruning, and then clusters the image anchors and the original image samples simultaneously, directly solving a variant of the original graph cut problem: the bipartite graph cut problem. To further simplify the optimization process, the above discrete model is transformed into the following matrix form:
[0064]
[0065] where Ind is a set of binary label matrices with only one element equal to 1 in each row, F and G are the discrete image sample label matrix and the image anchor label matrix respectively, and are the matrix representation forms of f ik and g jk respectively. However, the above matrix form has a trivial solution where all image samples and anchors belong to one class. At this time, the objective function is the sum of all elements of B and reaches the maximum value. Therefore, orthogonal normalization is performed on F and G to obtain the final model proposed by the present invention:
[0066]
[0067] where the introduced orthogonal normalization matrix restricts that there cannot be all-zero columns in the F and G matrices by the basic requirements of matrix inversion, thus effectively solving the above trivial solution.
[0068] where the discrete image sample label matrix and the image anchor label matrix are constructed as:
[0069] Define f jkis the element at the j-th row and k-th column of matrix F, and g jk is the element at the j-th row and k-th column of matrix G, and c is the number of categories included in the image dataset. Among them, the construction rules of matrices F and G are as follows: if the i-th image sample belongs to the k-th category, then f ik = 1, otherwise f ik = 0, to represent whether x i belongs to the k-th category. If the j-th image anchor belongs to the k-th category, then g jk = 1, otherwise g ik = 0, to represent whether z j belongs to the k-th category.
[0070] Then, randomly initialize F and G respectively according to the above construction rules;
[0071] Step 4: Optimize the final model by alternately iteratively updating F and G:
[0072] 1), Fix F and update G:
[0073] When F is fixed, the equivalence of the final model is:
[0074]
[0075] Among them, the constant matrix For the binarized label matrix G ∈ Ind, the coordinate ascent algorithm is used to update the label index row by row, that is, when updating one row, the other rows are fixed. Therefore, the above equivalence is transformed into the following vector form:
[0076]
[0077] Among them, v k and g k are the k-th column vectors of V and G respectively. Assume that the incremental matrix corresponding to G is For the fixed j-th anchor label row vector g j , t jk represents the increment of the objective function in the above vector form caused by the change of the k-th category label index of the j-th image anchor. Therefore, according to the construction of the objective function in vector form, the increment t jk corresponding to the j-th anchor and the k-th cluster is expressed as follows:
[0078]
[0079] Among them, <·> is a logical indicator, with a value of 1 if the logic is true, otherwise 0. is the binarized label row vector after the update of the j-th anchor.
[0080] Subsequently, traverse all m image anchors to complete the update of G.
[0081] 2) Fix G and update F:
[0082] When G is fixed, the final model is equivalent to:
[0083]
[0084] where the constant matrix
[0085] Transform the above formula into the following vector form:
[0086]
[0087] where u k and f k are the k-th column vectors of U and F respectively. Similarly, assume the incremental matrix corresponding to F is For the fixed i-th sample label row vector f i , represents the increment of the objective function in the above vector form caused by the change of the k-th class related label index of the i-th image sample. Similarly, the increment of the i-th image sample corresponding to the k-th cluster is expressed as follows:
[0088]
[0089] Based on the above incremental expression, update the image sample labels according to the maximum increment:
[0090]
[0091] where is the binarized label vector of the i-th sample after update.
[0092] Subsequently, traverse all n original samples to complete the update of F.
[0093] 3) Calculate the objective function value of the final model and judge whether it converges by the difference from the previous objective function value:
[0094] If it does not converge, return to 1); if it converges, go to step 5;
[0095] Step 5: Due to the probability binarization characteristic of F, the labels of all image samples can be directly obtained through F finally.
[0096] When the preprocessing process is completed, large-scale image data can be quickly retrieved by using the full-sample clustering center comparison or image anchor comparison.
[0097] There are two strategies for retrieving out-of-sample new image data: 1) Calculate the clustering center of each cluster based on the labels of all original image data samples. When introducing new image data, first compare it with c clustering centers, assign it to the cluster of the nearest clustering center, and then further search for the most similar original image sample in the corresponding cluster; 2) First use G to calculate the labels of all anchor images, and then directly compare the introduced new images with all anchors and assign the corresponding cluster of the nearest anchor to achieve retrieval.
[0098] Considering the unsupervised nature of the clustering technology, in this paper, when the number of anchors m = 2 a = 2 11 = 2048 and the number of bipartite graph neighbors r is 30, the clustering performance of the present invention is verified through repeated experiments on the MNIST image dataset: Input the original image data into the model proposed by the present invention and conduct 20 repeated experiments, and select the predicted label of the original sample corresponding to the maximum objective function among the 20 results. The clustering accuracy and normalized mutual information reach 69.10% and 70.90% respectively, and the single preprocessing time is 23.14 seconds. The variation of its objective function value and clustering accuracy with the number of iterations is as shown at the end of this paper Figure 2 As shown, it effectively verifies the convergence characteristics of the method proposed by the present invention; compared with the bipartite graph fast spectral clustering method proposed by Luo Xinglong et al., which only reaches 66.42% and 67.79% under the same parameter combination, it improves by 2.64 and 3.11 percentage points respectively, and significantly reduces the running time (227.71 seconds → 23.14 seconds). Thus, it can be seen that the method proposed by the present invention not only greatly improves the clustering efficiency, but also further improves the clustering accuracy, verifying its high efficiency and superiority.
[0099] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should be covered within the protection scope of the present invention.
Claims
1. A fast large-scale image data preprocessing method based on bipartite graph cut, characterized in that The steps are as follows: Step 1: Stretch and integrate n images of size a×b into an image data matrix where d = a×b is the number of pixels of a single image data; Step 2: Perform downsampling on the image data matrix X, that is, select m = 2 a anchor points using the hierarchical binary K-means algorithm, and obtain the anchor point data matrix where: a is the number of layers of the binary tree of the hierarchical binary K-means algorithm; Construct the following bipartite graph problem: Among them, b i is the i-th row vector of the bipartite graph matrix B, representing the i-th image vector x i and the bipartite graph similarity with all other image vectors, 1 m is a column vector with all m elements being 1; b ij is the element in the i-th row and j-th column of B, representing the i-th image vector x i and the bipartite graph similarity with the j-th anchor vector z j ; γ is the regularization parameter of the sparse regularization term; Solve the above bipartite graph problem by the following formula, and calculate the bipartite graph matrix with the number of nearest neighbors being r where r is the number of anchor neighbors for each sample; Step 3: Construct the discrete image sample label matrix and the image anchor label matrix respectively and define f as the element in the j-th row and k-th column of matrix F, g jk as the element in the j-th row and k-th column of matrix G, and c as the number of categories included in the image dataset; based on matrices F, G, and B, construct the following discrete model that maximizes the within-class similarity of the anchor layer and the sample layer: jk At the same time, perform orthogonal normalization on both F and G to obtain the discrete normalization model as follows: Then, randomly initialize F and G respectively according to the above construction rules; Step 4: Optimize the discrete normalization model constructed in Step 3 by alternately iteratively updating F and G: 1) Fix F and update G: Calculation and calculate the incremental matrix T row by row according to the following formula: Among them, v k and g k are the k-th column vectors of V and G respectively, and v jk and g jk are the j-th elements of the vectors v k and g k respectively, and t jk is the element in the j-th row and k-th column of the increment matrix T, representing the increment of the objective function corresponding to the change of the k-th class label index of the j-th image anchor point; Then, based on the incremental matrix T, update the indicator matrix G according to the following formula: where <·> is a logical indicator, with a value of 1 if the logic is true and 0 otherwise; is the updated label element of the j-th image anchor with respect to the k-th class; 2) Fix G and update F: Calculate and calculate the incremental matrix T row by row using the following formula * :[[]]END]] where, u k and f k are the k-th column vectors of U and F respectively, and u ik and f ik are the i-th elements of vectors u k and f k respectively; is the element at the i-th row and k-th column of the increment matrix T * , representing the increment of the objective function corresponding to the change of the k-th class label index of the i-th image sample point; Then, based on the incremental matrix T * the indication matrix F is updated according to the following formula: Among them, is the updated label element of the i-th image sample point with respect to the k-th class; 3) Calculate the objective function value of the discrete normalization model, and judge whether it converges by the difference from the previous objective function value: If it does not converge, return to ①; if it converges, proceed to Step 5; Step 5: Output F and G to obtain the predicted labels of n original image samples and m image anchors respectively. At this time, the image data preprocessing process ends.
2. The fast large-scale image data preprocessing method based on bipartite graph cut according to claim 1, wherein: The regularization parameter γ can be adaptively determined during the solution process of the bipartite graph problem.
3. The fast large-scale image data preprocessing method based on bipartite graph cut according to claim 1, wherein: The construction rules of the matrices F and G are as follows: if the i-th image sample belongs to the k-th class, then f ik = 1; otherwise, f ik = 0, to represent whether x i belongs to the k-th class. If the j-th image anchor belongs to the k-th class, then g jk = 1; otherwise, g ik = 0, to represent whether z j belongs to the k-th class.
4. Application of the fast large-scale image data preprocessing method based on bipartite graph cut according to any one of claims 1 to 3, characterized in that: After the preprocessing process is completed, full-sample clustering center comparison or image anchor comparison can be used to implement new image data retrieval outside the large-scale samples.
5. Application of the fast large-scale image data preprocessing method based on bipartite graph cut according to claim 4, characterized in that: The strategy for new image data retrieval outside the samples: 1) Calculate the clustering center of each cluster according to the labels of all original image data samples. When introducing new image data, first compare it with c clustering centers, assign it to the cluster of the nearest clustering center, and then further find the most similar original image sample in the corresponding cluster; 2) First use G to calculate the labels of all anchor images, and then directly compare the introduced new image with all anchors and assign the corresponding cluster of the nearest anchor to achieve retrieval.
Citation Information
Patent Citations
Rapid hyperspectral image clustering method and device, equipment and medium
CN111753904A
Direct and fast image clustering method based on bipartite graph collaborative clustering
CN115661496A