Semi-supervised image classification method based on dual label propagation
Through the semi-supervised learning method based on dual label propagation, the alternating optimization of the anchor matrix and the two-part graph affinity matrix is used to solve the problems of high time complexity and difficulty in identifying new categories in the existing methods, and high-precision image classification and new category discovery are achieved, which enhances the adaptability and robustness to real-world data.
Patent Information
- Application Number
- CN202510273448.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-07-08
AI Technical Summary
The existing semi-supervised learning methods based on graphs have problems such as high time complexity, high spatial complexity and inability to identify potential categories during the label propagation process. Especially when the prior label information is incomplete in real-world data, traditional methods cannot effectively identify new categories.
By constructing a semi-supervised learning model based on dual label propagation, the anchor matrix is generated using the clustering method, the two-part graph affinity matrix is constructed, and the two-part graph matrix and soft label matrix are updated through alternate optimization, and prior information constraints of anchor points and new class indicator items are introduced to realize the two-way transmission of label information and the discovery of new categories.
It improves the classification accuracy and robustness of semi-supervised learning, can effectively discover potential categories, reduces the time complexity of the algorithm, and improves the robustness and adaptability to noise labels.
Smart Images

Figure CN120279299A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of image processing technologies, and in particular, to a semi-supervised image classification method based on dual label propagation. Background Art
[0002] According to different supervision modes, machine learning is generally divided into categories such as supervised learning, unsupervised learning, and semi-supervised learning. Among them, supervised learning requires that all training samples carry prior label information, while unsupervised learning does not require any prior label information. Generally, the performance of supervised learning is better than that of unsupervised learning. However, the manual annotation of samples is extremely costly. Therefore, researchers have proposed semi-supervised learning, which uses a small number of labeled samples and a large number of unlabeled samples to train algorithms, reducing the cost of manual annotation and improving the algorithm performance as much as possible. The graph-based semi-supervised learning method mines potential structural information by establishing a similarity graph between all samples (including labeled samples and unlabeled samples), and then performs label propagation based on the similarity graph to propagate label information from labeled samples to unlabeled samples. However, in the real world, prior label information often cannot cover all data categories, and unlabeled samples may belong to unknown categories. Traditional graph-based semi-supervised learning methods cannot identify potential categories.
[0003] In recent years, graph-based semi-supervised learning methods with new class discovery capabilities have received extensive attention from scholars. Nie et al. (A general graph-based semi-supervised learning with novel class discovery. Neural Comput & Applic, 2010, 19:549–555.) proposed a method to introduce a class indicator for labeling new classes in graph-based semi-supervised methods through the General Graph-based Semi-Supervised Learning (GGSSL) method, thereby realizing the identification of potential classes. However, this method uses a fixed similarity graph for label propagation and often fails to obtain the optimal solution. Yuan et al. (A semi-supervised learning algorithm via adaptive Laplacian graph, Neurocomputing, 2021, 426: 162 - 173.) proposed the Adaptive Laplacian Graph Semi-Supervised Learning (ALGSSL) method, which introduces a label smoothing term and iteratively constructs the Laplacian graph and propagates labels to learn the optimal graph structure to improve semi-supervised learning performance. However, since this method requires multiple constructions of the Laplacian graph and needs to perform large-scale matrix inversion operations, it has high time complexity and space complexity.
[0004] Therefore, it is necessary to improve one or more problems existing in the above-mentioned related technical solutions.
[0005] It should be noted that this part aims to provide background or context for the technical solutions of the present disclosure stated in the claims. The description here is not admitted to be prior art just because it is included in this part. Summary of the Invention
[0006] The purpose of the embodiments of the present disclosure is to provide a semi-supervised image classification method based on dual label propagation, thereby at least to some extent overcoming one or more problems caused by the limitations and defects of related technologies.
[0007] According to the embodiments of the present disclosure, a semi-supervised image classification method based on dual label propagation is provided. The method includes: Convert the images in the image dataset into an image data matrix, and use a clustering method to generate m anchor points to form an anchor point matrix; Based on the image data matrix and the anchor matrix, an adaptive neighborhood assignment method is used to construct a bipartite graph affinity matrix between samples and anchors; Based on the bipartite graph affinity matrix, an objective function of a semi-supervised learning model based on dual label propagation is constructed; among them, the objective function includes the bipartite graph affinity matrix and the soft label matrix; By alternately optimizing and updating the bipartite graph matrix and the soft label matrix, the bipartite graph matrix and the soft label matrix with optimal structure are obtained; among them, the soft label matrix includes the sample label matrix and the anchor label matrix; According to the optimized soft label matrix, the class membership of samples and anchors is determined. When the class label exceeds the known class range, it is determined as an unknown new class.
[0008] Furthermore, in the step of converting the images in the image dataset into an image data matrix and generating m anchors by using a clustering method to form an anchor matrix, it includes: The one containing number of The image dataset of pixel-scale images is stretched into an image data matrix ; among them, each row of the image data matrix is a sample, is the number of images, is the total number of pixels of a single image; Using the Kmeans clustering method to generate representative anchors to obtain the anchor matrix ; among them, the number of anchors m is a preset parameter.
[0009] Furthermore, the expression of the bipartite graph affinity matrix is:
[0010]
[0011] Among them, is the anchor closest to the sample point the th is the sample point and the anchor the square of the Euclidean distance between them, is the preset number of nearest neighbors, is the sample point and the anchor the square of the Euclidean distance between them, is the sample point and the anchor the square of the Euclidean distance between them.
[0012] Furthermore, the expression of the objective function is:
[0013] Among them, is the similarity matrix, is the soft label matrix, is the prior label matrix, is the degree matrix, is the sum of the number of sample points and anchor points, is the finally obtained number of clusters, is the first weight, is the second weight, is the matrix trace operation, is the Laplacian graph, is the parameter matrix.
[0014] Furthermore, in the step of obtaining the bipartite graph matrix and the soft label matrix with optimal structure by alternately optimizing and updating the bipartite graph matrix and the soft label matrix, it includes: Fix the soft label matrix and update the bipartite graph matrix through a closed-form solution; Fix the bipartite graph matrix and update the soft label matrix through matrix derivation to obtain the bipartite graph matrix with optimal structure , the sample label matrix and the anchor label matrix .
[0015] Furthermore, when updating the bipartite graph matrix, solve the optimization problem with row sum constraints by the method of Lagrange multipliers, and the closed-form solution is:
[0016] Among them, , .
[0017] Furthermore, when updating the soft label matrix, obtain it by solving a linear equation:
[0018] Among them, is the sample label matrix, is the anchor label matrix, , is the identity matrix, is the parameter matrix The submatrix formed by the first n rows and columns, is the parameter matrix The submatrix formed by the last m rows and columns, is a diagonal matrix, and the diagonal elements are the column sums of the matrix , is the sample prior label matrix, is the anchor prior label matrix.
[0019] Further, in the step of determining the class attribution of samples and anchors according to the optimized soft label matrix and determining them as unknown new classes when the class labels exceed the known class range, the following steps are included: Let the sum of each row of the sample label matrix and the anchor label matrix be 1, then the final cluster classes to which the samples and anchors belong are:
[0020] where c is the class value of the samples and anchors, is the cluster class to which the sample belongs, is the cluster class to which the anchor belongs; When the class value belongs to , the samples and anchors belong to known classes; When the class value is , the samples and anchors belong to unknown new classes.
[0021] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: In the embodiments of the present disclosure, through the above semi-supervised image classification method based on dual label propagation, on the one hand, a dual label propagation framework is established. Innovatively, prior information constraints on the anchors are introduced, and the anchors are added as the starting points of label propagation, strengthening the influence factor of the anchors on the label transfer of the sample points during the label propagation process, realizing the two-way transfer of prior label information and learning results on the bipartite graph, thereby fully mining the sample point and anchor information stored in the rows and columns of the bipartite graph matrix, and improving the performance of the graph-based semi-supervised learning method; On the other hand, in this application, by alternately executing the bipartite graph construction and dual label propagation processes, a similarity relationship between samples and anchors is jointly established based on the distance measures in the original sample space and the low-dimensional label space, thereby constructing an optimal bipartite graph based on the clustering hypothesis and label smoothing hypothesis, avoiding the problem of performance degradation caused by the independent composition and label learning processes, and improving the classification accuracy of graph-based semi-supervised learning; On the third hand, considering the potential new classes in the data, this application introduces a new class indicator term to indicate the probability that the samples and anchors belong to new classes, and balances the prior information fitting term and the new class discovery term through a regularization parameter, thereby weakening the influence of possible wrong labels in the prior labels on the classification performance, discovering potential class information not included in the prior labels, and enhancing the adaptability and robustness of graph-based semi-supervised learning to real-world data. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present disclosure, and together with the specification are used to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0023] Figure 1 A step diagram showing a semi-supervised image classification method based on dual label propagation in an exemplary embodiment of the present disclosure; Figure 2 A specific flowchart showing a semi-supervised image classification method based on dual label propagation in an exemplary embodiment of the present disclosure; Figure 3 A line graph showing experimental results on the object image dataset Coil20 in an exemplary embodiment of the present disclosure. Detailed implementation manners
[0024] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described may be combined in any suitable manner in one or more embodiments.
[0025] In addition, the accompanying drawings are only schematic illustrations of the embodiments of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities.
[0026] In this example embodiment, a semi-supervised image classification method based on dual label propagation is provided. As shown in Figure 1 , the semi-supervised image classification method based on dual label propagation may include: step S101 to step S105.
[0027] Step S101: Convert the images in the image dataset into an image data matrix, and generate m anchor points using a clustering method to form an anchor point matrix; Step S102: Based on the image data matrix and the anchor point matrix, construct a bipartite graph affinity matrix between the samples and the anchor points using an adaptive neighborhood assignment method; Step S103: Based on the bipartite graph affinity matrix, construct an objective function of a semi-supervised learning model based on dual label propagation; wherein, the objective function includes the bipartite graph affinity matrix and a soft label matrix; Step S104: Update the bipartite graph matrix and the soft label matrix through alternating optimization to obtain the bipartite graph matrix and the soft label matrix with optimal structure; wherein, the soft label matrix includes a sample label matrix and an anchor label matrix. Step S105: Determine the class membership of samples and anchors according to the optimized soft label matrix. When the class label exceeds the known class range, it is determined as an unknown new class.
[0028] Through the above semi-supervised image classification method based on dual label propagation, on the one hand, a dual label propagation framework is established. The prior information constraint of the anchor is innovatively introduced, and the anchor is added as the starting point of label propagation, strengthening the influence factor of the anchor on the label transfer of sample points during the label propagation process, realizing the bidirectional transfer of prior label information and learning results on the bipartite graph, thereby fully mining the sample points and anchor information stored in the rows and columns of the bipartite graph matrix, and improving the performance of the graph-based semi-supervised learning method. On the other hand, in this application, by alternately executing the bipartite graph construction and dual label propagation processes, the similarity relationship between samples and anchors is jointly established based on the distance measures in the original sample space and the low-dimensional label space, thereby constructing the bipartite graph with optimal structure based on the clustering hypothesis and the label smoothing hypothesis, avoiding the performance degradation problem caused by the independence of the graph construction and label learning processes, and improving the classification accuracy of the graph-based semi-supervised learning. On the third hand, considering the potential new classes in the data, this application introduces a new class indicator term to indicate the probability possibility that samples and anchors belong to new classes, and balances the prior information fitting term and the new class discovery term through a regularization parameter, thereby weakening the influence of the possible mislabels in the prior labels on the classification performance, discovering the potential class information not included in the prior labels, and enhancing the adaptability and robustness of the graph-based semi-supervised learning to real-world data.
[0029] Next, reference will be made to Figures 1 to 3 to describe each step of the above semi-supervised image classification method based on dual label propagation in this exemplary embodiment in more detail.
[0030] In step S101, the images in the image dataset are converted into an image data matrix, and m anchors are generated by using a clustering method to form an anchor matrix.
[0031] Specifically, for the image dataset, the dataset containing images with a pixel scale of is stretched into an image data matrix , each row of which is a sample, is the number of images, is the total number of pixels of a single image, that is, the feature dimension of the image. The Kmeans clustering method is used to generate representative anchors to obtain the anchor matrix , where is the number of artificially selected anchor points.
[0032] In step S102, based on the image data matrix and the anchor point matrix, an adaptive neighborhood assignment method is used to construct a bipartite graph affinity matrix between samples and anchor points.
[0033] Specifically, based on the image data matrix of the previous step and the obtained anchor point matrix , an adaptive neighborhood assignment method is used to construct a bipartite graph affinity matrix . The result obtained by the adaptive neighborhood assignment method is as follows: as follows:
[0034] where represents the anchor point closest to the sample point the th represents the sample point and the anchor point the square of the Euclidean distance between them, is an artificially selected parameter.
[0035] In step S103, based on the bipartite graph affinity matrix, an objective function of a semi-supervised learning model based on dual label propagation is constructed; where the objective function includes the bipartite graph affinity matrix and the soft label matrix.
[0036] Specifically, (3) based on the constructed bipartite graph affinity matrix , the objective function of the semi-supervised learning model based on dual label propagation is constructed as follows:
[0037] where is the similarity matrix, is the soft label matrix, is the prior label matrix, is the degree matrix; is the sum of the number of sample points and anchor points, is the number of clusters finally obtained. In the above formula, the first two terms are the adaptive neighborhood bipartite graph construction terms, and the last two terms are the dual label propagation terms. The prior label matrix can be split into: , where is the sample prior label information. If the labeled sample belongs to the th class, then , and the other elements in this row are 0; for the unlabeled sample , the other elements of this row are 0. Similarly, is the anchor prior label information. For all anchors , the other elements are 0. In the above formula, the similarity matrix is: , Thus, the above model can be further transformed into:
[0038] where is the Laplacian graph.
[0039] In step S104, by alternately optimizing and updating the bipartite graph matrix and the soft label matrix, the bipartite graph matrix and the soft label matrix with optimal structure are obtained; among them, the soft label matrix includes the sample label matrix and the anchor label matrix.
[0040] Specifically, the optimal solution of the above problem is obtained by alternately solving.
[0041] Bipartite graph update: Fix the label , and update the bipartite graph matrix
[0042] When is fixed, the optimization problem is equivalent to:
[0043] Expanding the above matrix trace problem, we can get:
[0044] Furthermore, the optimization problem can be written as:
[0045] In the above formula, updating each row of the matrix is independent of each other. Splitting the optimization objective by the rows of the matrix , we can get:
[0046] Applying the Lagrange multiplier method to solve this problem, its Lagrangian function is:
[0047] where .
[0048] The closed-form solution of this Lagrange multiplier problem is:
[0049] where , the parameters can be obtained from the constraints of the optimization problem matrix , and then the final solution is: , then the final solution is:
[0050] Label propagation: Fix the bipartite graph matrix , and update the labels
[0051] When is fixed, the optimization problem is equivalent to:
[0052] Taking the partial derivative of the above formula with respect to the matrix yields:
[0053] To distinguish between anchor points and sample points, introduce the parameter and rewrite the matrix as follows:
[0054] where is the zero matrix, is the identity matrix. Substituting the above formula into the partial derivative result, the label propagation process can be obtained as:
[0055] Let , then from the above formula, we can get:
[0056] In step S105, according to the optimized soft label matrix, determine the class membership of the samples and anchor points. When the class label exceeds the known class range, it is determined as an unknown new class.
[0057] Specifically, through alternating optimization, the structurally optimal bipartite graph matrix , the sample label matrix and the anchor point label matrix can be obtained. It should be noted that in this application, the label matrix always maintains a probabilistic meaning, that is, the sum of each row of the label matrix is 1. Then the final cluster classes to which the samples and anchor points belong are obtained by the following formula:
[0058] From the above formula, it can be seen that when the class value belongs to , the samples and anchor points belong to the known classes; when the class value is , the samples and anchor points belong to the unknown class, that is, the new class discovery function is realized.
[0059] In a specific embodiment, the present application proposes a semi-supervised image classification method based on dual label propagation. Taking the object image dataset Coil20 as an example, the specific steps of the proposed method for image classification are elaborated, but the technical content of the present application is not limited to the described scope. The object image dataset Coil20 contains a total of 1440 object images, each image has a length and width of 128 pixels, there are 20 categories in total, and each category has 72 sample points.
[0060] The present application proposes a semi-supervised image classification method based on dual label propagation, including the following steps, as Figure 2 shown, which is the specific flowchart of the semi-supervised image classification method based on dual label propagation.
[0061] 1. Downsample the images in the Coil20 dataset to images with a length and width of 32 pixels each. Use the grayscale features of the images as the data features of the images. Straighten the pixel grayscale values of each image into a vector, and the dimension of the vector is 1024 dimensions. The corresponding feature matrix of the image is , where each row of the matrix is a sample, is the number of image samples, is the dimension of the feature. Use the Kmeans clustering method to generate representative anchor points, generate anchor points, and obtain the anchor point matrix , where is the number of anchor points selected manually.
[0062] 2. Based on the image data matrix from the previous step and the obtained anchor point matrix , construct the bipartite graph affinity matrix :
[0063] Initialize the following parameters before iterative solution: (1) Calculate the matrix is a diagonal matrix, and its th diagonal element ; (2) The matrix , , so that the row sum is 1; (3) The matrix .
[0064] 3. Update the bipartite graph matrix , define as the number of iterations , where , update 。
[0065] 4. Update the label matrix , : , wherein 。
[0066] 5. Repeat to update the bipartite graph matrix and the label matrix and until the algorithm converges, output the sample label matrix and the anchor label matrix The column where the maximum value of each row of the label matrix is located represents the category of the sample point represented by that row, that is:
[0067] When it indicates that the sample is classified into a pre-known category; When it shows that the sample point belongs to a potential new category. Thus, the method proposed in this application has completed the image classification task on the object image dataset Coil20 and can indicate the image categories not in the prior information.
[0068] To verify the effectiveness and efficiency of the image classification method of this application in performing classification on the object image dataset Coil20, the average classification accuracy in five repeated experiments is selected as the final classification result. In this experiment, the value range of the number of anchor points is , the label ratio is 10%. Further, to analyze the ability of the algorithm to cope with noise interference, the label noise ratio is set to 0%, 5%, 10%, 15%, 20%. In this application, a priori information constraint is innovatively imposed on the anchor points. By iteratively performing bipartite graph construction and dual label propagation, and combining the label smoothing term to learn the optimal bipartite graph, the time complexity of the algorithm can be significantly reduced, the classification accuracy of the algorithm can be improved, and the robustness of the algorithm to noisy labels can be enhanced. According to Figure 3It can be seen from the experimental performance curve that the image classification accuracy of the algorithm of the present application is much higher than that of the GGSSL and ALGSSL algorithms. Moreover, when the proportion of noisy labels increases, the performance degradation trend of the algorithm of the present application is less than that of other comparison algorithms, thus verifying the effectiveness of the proposed dual label propagation method. In addition, to verify the performance of the present application in dealing with the problem of new class discovery, experiments are performed on the public image dataset Pendigits, which has 10 categories. In the experiment, 3 of the categories are set as unknown new classes, and the GGSSL, ALGSSL and the algorithm of the present application are respectively executed. The classification accuracy of the present application for the unknown new classes is 98.10%. Correspondingly, the accuracies of GGSSL and ALGSSL are 98.09% and 97.77% respectively. It can be seen that the present application can more accurately discover unknown new class data.
[0069] Through the above semi-supervised image classification method based on dual label propagation, on the first hand, a dual label propagation framework is established. The prior information constraint of the anchor points is innovatively introduced, and the anchor points are added as the starting point of label propagation, strengthening the influence factor of the anchor points on the label transfer of sample points during the label propagation process, realizing the bidirectional transfer of prior label information and learning results on the bipartite graph, thus fully mining the sample point and anchor point information stored in the rows and columns of the bipartite graph matrix and improving the performance of the graph-based semi-supervised learning method. On the second hand, the present application alternately executes the bipartite graph construction and the dual label propagation process, and collaboratively establishes the similarity relationship between samples and anchor points based on the distance measures in the original sample space and the low-dimensional label space, thus constructing the bipartite graph with the optimal structure based on the clustering hypothesis and the label smoothing hypothesis, avoiding the performance degradation problem caused by the independence of the graph construction and label learning processes, and improving the classification accuracy of the graph-based semi-supervised learning. On the third hand, considering the potential new classes in the data, the present application introduces a new class indicator term to indicate the probability possibility that samples and anchor points belong to new classes, and balances the prior information fitting term and the new class discovery term through the regularization parameter, thus weakening the influence of the possible wrong labels in the prior labels on the classification performance, discovering the potential class information not included in the prior labels, and enhancing the adaptability and robustness of the graph-based semi-supervised learning to real-world data.
[0070] It should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "length", "width", "thickness", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "clockwise", "counterclockwise", etc. in the above description is the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the embodiments of the present disclosure and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation, and therefore should not be construed as a limitation to the embodiments of the present disclosure.
[0071] In addition, the terms "first" and "second" are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present disclosure, the meaning of "a plurality" is two or more, unless otherwise specifically defined.
[0072] In the embodiments of the present disclosure, unless otherwise clearly specified and limited, the terms such as "mounted", "connected", "connected to", "fixed" and the like should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two elements or the interaction relationship between two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present disclosure can be understood according to specific circumstances.
[0073] In the embodiments of the present disclosure, unless otherwise clearly specified and limited, the first feature being "on" or "under" the second feature may include the direct contact between the first and second features, or may include the situation where the first and second features are not in direct contact but in contact through other features therebetween. Moreover, the first feature being "above", "over" and "on top of" the second feature includes that the first feature is directly above and obliquely above the second feature, or merely means that the horizontal height of the first feature is higher than that of the second feature. The first feature being "under", "below" and "beneath" the second feature includes that the first feature is directly below and obliquely below the second feature, or merely means that the horizontal height of the first feature is lower than that of the second feature.
[0074] In the description of this specification, the descriptions with reference to terms such as "one embodiment", "some embodiments", "example", "specific example" or "some examples" etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.
[0075] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the appended claims.
Claims
1. A semi-supervised image classification method based on dual label propagation, characterized in that The method includes: Converting the images in the image dataset into an image data matrix, and generating m anchor points by using a clustering method to form an anchor point matrix; Based on the image data matrix and the anchor point matrix, constructing a bipartite graph affinity matrix between samples and anchor points by using an adaptive neighborhood assignment method; Based on the bipartite graph affinity matrix, constructing an objective function of a semi-supervised learning model based on dual label propagation; wherein, the objective function includes the bipartite graph affinity matrix and a soft label matrix; By alternately optimizing and updating the bipartite graph matrix and the soft label matrix, obtaining a bipartite graph matrix and a soft label matrix with optimal structures; wherein, the soft label matrix includes a sample label matrix and an anchor point label matrix; According to the optimized soft label matrix, determining the class membership of samples and anchor points, and when the class label exceeds the known class range, determining it as an unknown new class.
2. The semi-supervised image classification method based on dual label propagation according to claim 1, characterized in that In the step of converting the images in the image dataset into an image data matrix and generating m anchor points by using a clustering method to form an anchor point matrix, it includes: Include pieces of Stretch the image data set containing an image of pixel scale into an image data matrix ; Wherein, each row of the image data matrix is a sample, is the number of images, is the total number of pixels of a single image; Generate using the Kmeans clustering method representative anchor points to obtain an anchor point matrix ; where the number of anchor points m is a preset parameter.
3. The semi-supervised image classification method based on dual label propagation according to claim 2, characterized in that The expression of the bipartite graph affinity matrix is: Among them, is the anchor point closest to the sample point the nearest one, is the square of the Euclidean distance between the sample point and the anchor point ; is the preset number of nearest neighbors, is the square of the Euclidean distance between the sample point and the anchor point ; is the square of the Euclidean distance between the sample point and the anchor point .
4. The semi-supervised image classification method based on dual label propagation according to claim 3, wherein The expression of the objective function is: Among them, is the similarity matrix, is the soft label matrix, is the prior label matrix, is the degree matrix, is the sum of the number of sample points and anchor points, is the finally obtained number of clusters, is the first weight, is the second weight, is the matrix trace operation, is the Laplacian graph, is the parameter matrix.
5. The semi-supervised image classification method based on dual label propagation according to claim 4, characterized in that In the step of obtaining a bipartite graph matrix and a soft label matrix with optimal structures by alternately optimizing and updating the bipartite graph matrix and the soft label matrix, it includes: Fixing the soft label matrix and updating the bipartite graph matrix through a closed-form solution; Fix the bipartite graph matrix, update the soft label matrix through matrix differentiation to obtain the bipartite graph matrix with optimal structure , the sample label matrix and the anchor label matrix .
6. The semi-supervised image classification method based on dual label propagation according to claim 5, characterized in that, When updating the bipartite graph matrix, solving the optimization problem with row sum constraints by using the Lagrange multiplier method, and the obtained closed-form solution is: Among them, , .
7. The semi-supervised image classification method based on dual label propagation according to claim 6, wherein When updating the soft label matrix, obtaining it by solving a linear equation: Among them, is the sample label matrix, is the anchor label matrix, , is the identity matrix, is the parameter matrix The sub-matrix formed by the first n rows and columns, is the parameter matrix The sub-matrix formed by the last m rows and columns, is a diagonal matrix, and the diagonal elements are the column sums of the matrix , is the sample prior label matrix, is the anchor prior label matrix.
8. The semi-supervised image classification method based on dual label propagation according to claim 7, wherein In the step of determining the class membership of samples and anchor points according to the optimized soft label matrix and determining it as an unknown new class when the class label exceeds the known class range, it includes: Let the sum of each row of the sample label matrix and the anchor label matrix be 1, then the cluster categories to which the final samples and anchors belong are: where c is the class value of the sample and the anchor point, is the cluster class to which the sample belongs, is the cluster class to which the anchor point belongs; When the category value belongs to the sample and the anchor point belong to the known category; When the class value is the sample and the anchor belong to an unknown new class.