Online image classification method based on mark distribution learning in open environment
By adopting an online image classification method based on label distribution learning in an open environment, dynamically adjusting the classification model and introducing a noise filtering module, the problem of traditional methods being difficult to cope with dynamic changes in categories and uneven data distribution is solved, and the robustness and classification accuracy of the model are significantly improved.
Patent Information
- Application Number
- CN202510317329.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-20
AI Technical Summary
Traditional image classification methods are difficult to deal with the problems of dynamic changes in categories, labeling noise, and uneven data distribution in open environments, especially when new categories appear and old categories disappear.
The online image classification method based on label distribution learning is adopted, and the classification model is dynamically adjusted, and the noise filtering module is introduced to improve the robustness and classification accuracy of the model.
This method can effectively identify and filter noise marks, improve the robustness and classification accuracy of the model, and adapt to the problems of dynamic categories and uneven data distribution in open environments.
Smart Images

Figure CN120182707A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of computer vision and machine learning, and particularly to an online image classification method based on label distribution learning in an open environment. Background Art
[0002] With the explosive growth of image data, image classification has become an important application in fields such as autonomous driving, intelligent monitoring, and industrial inspection. However, traditional image classification methods usually assume that the categories are fixed and known, and cannot effectively cope with problems such as dynamically changing categories, label noise, and unbalanced data distribution in an open environment. In an open environment (such as dynamic scenes, real-time data streams, changing classification labels, etc.), traditional image classification methods face the following challenges: 1. The continuous emergence of new categories and the disappearance of old categories.
[0003] 2. The data distribution in an open environment changes over time, and traditional static classification models are difficult to adapt to long-term changes.
[0004] 3. Online classification requires the model to be updated quickly and complete the classification task within a limited time.
[0005] 4. Unbalanced data distribution causes the model to be biased towards the majority categories.
[0006] Label Distribution Learning (LDL), as a new learning paradigm, can utilize label distributions to represent the diversity of samples.
[0007] However, current LDL-based methods are mainly used in static scenarios and lack research on online image classification in an open environment. Summary of the Invention
[0008] Each exemplary embodiment of this application provides an online image classification method based on label distribution learning in an open environment, so as to at least have the technical effects of identifying and filtering noise labels, and significantly improving the robustness and classification accuracy of the model.
[0009] According to one aspect of this application, each exemplary embodiment of this application provides an online image classification method based on label distribution learning in an open environment, and the method includes the following steps: An online image classification method based on label distribution learning in an open environment, the method includes the following steps: S1. Construct an image classification model based on label distribution learning, obtain the initial dataset, perform feature increment on the initial dataset, perform feature dimensionality reduction and feature selection on the received new features, enrich the features and compress the data to construct a label distribution matrix to represent the relationship between the image and the category, and obtain a new dataset after feature dimensionality reduction; S2. For the new dataset, update the model parameters online, mine the topological structure of the feature space, analyze the sample correlation through the sparse learning method, represent the target sample through the linear combination of training samples, capture the mapping from the feature space to the label space, and transfer it to the label space to make the label space also satisfy a similar sparse structure; S3. Design a category dynamic adjustment mechanism to automatically identify new categories and eliminate invalid categories; S4. Introduce a noise filtering module to reduce the impact of label noise on the model; S5. Retrain and evaluate the model with the filtered dataset, and adjust the parameters of the noise filtering module according to the evaluation results to finally obtain an optimized model.
[0010] The present application has the following beneficial effects: By introducing the label distribution learning mechanism, this method dynamically adjusts the classification model to adapt to the emergence of new categories and the disappearance of old categories, and at the same time reduces the impact of label noise on the model through the noise filtering module; First, aiming at the characteristics of streaming big data arriving dynamically, large sample quantity scale, and unknown sample dimension, a feature dimensionality reduction algorithm based on incremental kernel is designed. Secondly, for the situation of redundant and useless features existing in the dimensionality-reduced feature space, the present invention proposes a method based on the correlation of sparse representation samples, fully mining the topological structure information of the feature space and the correlation information between labels, providing more supervision information for subsequent label distribution learning prediction; Then, in order to fully learn label correlation, this project introduces an adaptive graph, considering the correlation between the label space and the feature space at the same time. In short, the feature subset obtained through a series of feature learning can make the subsequent incremental learning of the system more efficient; Finally, introduce a noise filtering module: analyze the outliers in the label distribution matrix; identify and filter the label noise; improve the robustness and classification accuracy of the model.
[0011] The noise filtering module can significantly improve the robustness and classification accuracy of the model by analyzing the outliers in the label distribution matrix, identifying and filtering the noise labels. Its core steps include the construction of the label distribution matrix, outlier detection, noise label identification, noise filtering, and model training and evaluation. By reasonably designing the noise filtering module, the label noise problem can be effectively addressed, and the performance of the model in practical applications can be improved. Description of the Drawings
[0012] The accompanying drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings: Figure 1 It is a schematic diagram of the overall process involved in the method of the present invention; Figure 2 It is a system framework diagram involved in the method of the present invention; Figure 3 It is a schematic diagram of the generation of the label distribution matrix involved in the method of the present invention; Figure 4 It is a flowchart of online updating model parameters involved in the method of the present invention; Figure 5 It is a schematic diagram of the category dynamic adjustment mechanism involved in the method of the present invention; Figure 6 It is a flowchart of the operation of the noise filtering module involved in the method of the present invention. Detailed implementation manners
[0013] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the preferred embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments.
[0014] All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts fall within the scope of protection of the present application.
[0015] As Figures 1 to 6 shown, the present invention provides an online image classification method based on label distribution learning in an open environment. The method includes the following steps: S1. Construct an image classification model based on label distribution learning, obtain an initial data set, perform feature increment on the initial data set, perform feature dimensionality reduction and feature selection on the received new features, enrich the features and compress the data to construct a label distribution matrix to represent the relationship between the image and the category, and obtain a new data set after feature dimensionality reduction; S2. For the new data set, update the model parameters online, mine the topological structure of the feature space, analyze the sample correlation through a sparse learning method, represent the specified sample through a linear combination of training samples, capture the mapping from the feature space to the label space, and transfer it to the label space so that the label space also satisfies a similar sparse structure; S3. Design a category dynamic adjustment mechanism to automatically identify new categories and eliminate invalid categories; S4. Introduce a noise filtering module to reduce the impact of label noise on the model; S5. Retrain and evaluate the model using the filtered dataset, and adjust the parameters of the noise filtering module according to the evaluation results to finally obtain an optimized model.
[0016] It should be noted that the labeled distribution matrix is generated from the joint probability distribution of image features and class labels and is used to represent the association strength between images and classes; online updating of model parameters adopts an incremental learning algorithm to gradually optimize model parameters by minimizing the loss function; the class dynamic adjustment mechanism automatically identifies new classes and eliminates invalid classes with confidence levels lower than a preset threshold by calculating class confidence levels; the noise filtering module identifies and filters labeled noise by analyzing outliers in the labeled distribution matrix.
[0017] In one embodiment, S1 specifically includes: S11. The initial dataset is a small-sample dataset, denoted as representing a q-dimensional sample space, representing the initial labeled set and the initial labeled distribution set, , and representing the labeled distribution of instance . The initial dataset includes labeled known data, unlabeled known data, and unlabeled unknown data, denoted as respectively, where is a q-dimensional feature vector, expressed as ; S12. Use a preset classifier to predict the labeled distribution of data and incorporate it into the dataset . Denote the projection vector as p, and the feature vector of the sample can be projected into a new feature vector through the function , expressed as the kernel function: ; S13. Map the labeled distribution through a simple linear kernel function, ; Define the incremental kernel matrix in the feature space, where ; The kernel matrix of the labeled distribution , where ; Here, the kernel function is replaced by the nuclear norm, and it is transformed into a convex optimization problem: (1) where is the i-th singular value of matrix K, and the nuclear norm is a convex function, so low-dimensional feature mapping can be found by solving the convex optimization; S14. Incremental HSIC in an open environment, , where tr(·) is the trace of the matrix, , e is a column vector of all 1s, F and G are reproducing kernel Hilbert spaces mapped from X and d respectively. Substituting into the above formula and removing the normalization term, the optimal solution can be obtained: (2) Denote , and , ; S15, after incremental HSIC, the new dataset after feature dimensionality reduction can be obtained , where .
[0018] In one embodiment, the S2 specifically includes: S21, find a suitable label distribution on the new dataset, so that its distance from is the smallest. is the sparse representation of the sample points in the feature space, and is obtained by finding the sparse representation in the feature matrix to mine the global structure of the feature space; S22, use the method of sparse representation to model the relationship between a single example and other examples, so as to obtain the sparse coefficients; S23, use the quasi-Newton method to solve the optimal parameters.
[0019] In one embodiment, the S3 specifically includes:
[0020] S3.1, obtain a sparse and clearly structured adaptive graph, and the number of connected components in the adaptive graph is equal to the number of datasets / groups. S3.2, learn the adaptive graph, impose structural constraints on the Laplacian graph, and introduce an adaptive structured graph for semi-supervised learning in an open environment; S3.3, combine with the maximum entropy model for optimization and update its parameters.
[0021] In one embodiment, the S4 specifically includes: S41, construct a label distribution matrix, input the training dataset , where is the sample, is the label; for each sample , calculate its association strength with all classes, generate a label distribution vector , where C is the number of classes; combine the label distribution vectors of all samples into a label distribution matrix ; S42. Calculate the statistical features of each labeled distribution vector, identify outliers using thresholds or statistical tests; cluster the labeled distribution matrix and label samples far from the cluster center as outliers; calculate the similarity between sample labeled distribution vectors and identify labeled distribution vectors that are significantly different from other samples as outliers. S43. Evaluate the reliability of labels using the confidence of model predictions, and identify labels with confidence lower than the threshold as noise; perform a consistency test using the results of multiple models or multiple trainings, and identify inconsistent labels as noise; analyze abnormal patterns in the labeled distribution matrix and identify labels that do not conform to the normal distribution as noise. S44. According to the model prediction results or the labeled distribution matrix, correct the noisy labels, replacing the noisy labels with the categories predicted by the model; assign weights to each sample, reduce the weights of noisy samples, and use confidence or consistency scores as weights.
[0022] The following combines the accompanying drawings and elaborates on the execution process and principle of the complete method of this application through a specific embodiment.
[0023] The specific implementation manner of the present invention is as follows:
[0024] When performing labeled distribution learning in an open environment, there is some initial training data, which is used to train an initial model, and the initial data set is a small sample data set. Denote as q the -dimensional sample space, as the initial label set, the initial labeled distribution set and denote the labeled distribution of instance . Given the initial small sample data set, this data set contains three types of data, namely labeled known data, unlabeled known data, and unlabeled unknown data, which are respectively denoted as , where q is a T -dimensional feature vector, expressed as
[0025] . The task of feature increment is to be able to perform feature dimensionality reduction and feature selection when receiving new features, compress data while enriching features, and accelerate the speed of the model. Predict the labeled distribution of data through a preset classifier and incorporate it into the data set . Now, assume the data set dimension), in this feature space, there is a maximized relationship between the features of the samples and their label distributions. Denote the projection vector as p , the feature vector of the sample can be projected into a new feature vector through the function . This project describes this process as a kernel function: . Similarly, the label distribution can be mapped through a simple linear kernel function . According to the above analysis, define the incremental kernel matrix of the feature space , where ; the kernel matrix of the label distribution , where . Here, the kernel function can be replaced by the nuclear norm, converting it into a convex optimization problem.
[0026] (1) where is the i -th singular value of the matrix K, and the nuclear norm is a convex function. Therefore, a low-dimensional feature mapping can be found by solving the convex optimization. According to HSIC, , where tr(·) is the trace of the matrix, , e is a column vector of all 1s, F , G are the reproducing kernel Hilbert spaces mapped by X , d respectively. Substitute into the above formula and remove the normalization term to obtain the optimal solution: (2) Denote , and , .
[0027] The above operations are called incremental HSIC in an open environment. After incremental HSIC, a new dataset after feature dimensionality reduction can be obtained, where .
[0028] (2) Sample correlation learning Based on the new dataset, this project uses a feature selection method based on the label distribution learning model to select features that are more relevant to the label distribution. This can not only retain the feature classification ability of the dataset, but also effectively reduce the complexity of label distribution learning. In the output model of label distribution learning, the mapping relationship between x and y is represented by the parameter . Therefore, the instancex The predicted label y can be mapped using , similar to existing multi-label feature ranking algorithms, where the parameter reflects the importance of each feature. By means of the parameter , this project can obtain the significance of each feature by optimizing the objective function. Therefore, the importance of each feature can be ranked. From the perspective of feature selection, some features are redundant or negligible in classification, and not all features are useful for classification, because for each label, there are inevitably some redundant features. Given a matrix W , assuming the element W ij represents i- the importance of the feature for j -label. This project intends to use sparse learning methods to improve the effectiveness of multi-label feature selection methods. This project uses sparse representation to analyze sample correlation, especially the topological structure information in the feature space. By linearly combining training samples to represent a specified sample, the sample correlation between training examples is then analyzed. Sparse representation is used to mine the sample correlation information in the feature space and transfer it to the label space, so that the label space also satisfies a similar sparse structure. That is to say, this project hopes to find a suitable label distribution on the new dataset, making its distance from as small as possible. is the sparse representation of the sample points in the feature space, which is obtained by finding the sparse representation in the feature matrix to mine the global structure of the feature space. Then the function for mining the topological structure information of the feature space can be expressed as follows: (3) To mine the sample correlation in the feature space, this project uses the method of sparse representation to model the relationship between a single example and other examples, so as to obtain sparse coefficients, and mines the sample correlation in this globally refined way. To improve the effectiveness of feature selection methods, the purpose of feature selection technology is to select a strong discriminative feature subset to improve the performance of the learning model. The most important aspect is that the learning model can capture the mapping from the feature space to the label space. In most existing label distribution learning algorithms, the Kullback-Leibler (KL) divergence is used to measure the loss value of the basic objective function. This project combines the maximum entropy model provided by the KL divergence with sparse learning, which can accommodate the meaning of features for all labels. And the objective function is defined as: (4) To solve for the optimal parameter , this project uses the quasi-Newton method to solve. Compared with the Newton method, it can avoid the complex process of directly calculating the inverse matrix of the Hessian matrix.
[0029] (3) Label correlation learning As mentioned above, in the field of labeled distribution learning, it is difficult to perform feature selection tasks based on probability values such as labeled distribution because the labeled distribution cannot define the contribution of the feature set to the sample label. In addition, traditional feature dimensionality reduction methods are applied in closed learning environments. For new sample features, traditional methods will not be able to adapt to the incremental open environment, resulting in a dimensional mismatch between the sample feature set and the new sample features. How to adapt to the incremental open environment, continuously absorb new features and timely incorporate them into the low-dimensional feature set will become a very challenging problem. In order to adapt to the incremental environment, traditional semi-supervised learning (SSL) methods can effectively use large repositories of unlabeled data to improve performance while relying on a small set of labeled data. A common assumption in most SSL methods is that labeled data and unlabeled data come from the same data distribution. However, in many real-world scenarios, this is not the case, which limits their applicability. This project attempts to solve challenging open environment problems. To solve these problems, this project considers how to learn an adaptive graph that best fits the sample correlations in the label space and feature space. The adaptive graph should be sparse, with a clear structure, and the number of connected components in the graph is exactly the number of data clusters / classes. Such a structured graph will benefit many subsequent tasks as it contains more accurate information on data dependencies.
[0030] Specifically, through incremental dimensionality reduction and sample correlation mining, this project generates pseudo labels that can be treated isomorphically with ground truth labels. This makes it possible to use a unified maximum entropy loss on both labeled and unlabeled sets. In more detail, given a batch of images, feature similarity can provide guidance for mining hidden relationships in feature space and label space. For example, according to the smoothness hypothesis, the topological structure of the feature space can be transferred locally to the local digital annotation space, so these instance features can represent the similarity level of the labels. This project considers the correlation between the label space and the feature space at the same time, and then allows the topological information of the feature space to more fully guide the labeling of relevant information by incorporating adaptive graph learning into the objective function. In order to learn an adaptive graph with a clear structure, structural constraints are imposed on the Laplacian graph. This is the first time that an adaptive structured graph has been introduced for open environment semi-supervised learning. The maximum entropy model is then combined for optimization and its parameters are updated. For the instance space represents the complete tag set, for the instances in U , with its labeled distribution This project is aimed at For any instance in the label space L, a brand-new similarity is defined , which simultaneously considers the correlation between the label space and the feature space: (5) where represents the similarity matrix of the feature space.
[0031] Considering the probability of nearest neighbors to learn the similarity matrix, the probability that two data points are adjacent can be regarded as their similarity. Naturally, it is considered that the smaller the distance, the greater the probability of label similarity, and vice versa. For simplicity, the Euclidean distance is adopted here. Therefore, this project can adaptively determine the probability by solving the following problem (6) , and are three trade-off parameters. Among them, is the graph Laplacian matrix, is a diagonal matrix, and the diagonal elements are composed of . Substituting it in, the final optimization objective function T (manifold regularization term with semi-supervised adaptive graph): (7) After solving for the optimal parameters, use to evaluate the importance of features, and return the feature set in descending order . This project can select the top several features as the new feature subset of the instance .
[0033] (4) Noise filtering module
[0034] Input: Training data set , where is the sample, is the label.
[0035] Construct the label distribution matrix: For each sample , calculate its association strength with all classes to generate the label distribution vector , where C is the number of classes. Combine the label distribution vectors of all samples into the label distribution matrix .
[0036] Calculate the statistical features (such as mean, variance, skewness, etc.) of each token distribution vector, and identify outliers using thresholds or statistical tests (such as Grubbs test, Z-score). Cluster the token distribution matrix (such as K-means, DBSCAN), and mark the samples far from the cluster center as outliers. Calculate the similarity between the sample token distribution vectors (such as cosine similarity, Euclidean distance), and identify the token distribution vectors with large differences from other samples as outliers.
[0037] Evaluate the reliability of tokens using the confidence of model prediction (such as Softmax output), and identify the tokens with confidence lower than the threshold as noise. Conduct a consistency test using the results of multiple models or multiple trainings, and identify the inconsistent tokens as noise. Analyze the abnormal patterns in the token distribution matrix (such as the token distribution of certain categories being too concentrated or dispersed), and identify the tokens that do not conform to the normal distribution as noise.
[0038] According to the model prediction results or the token distribution matrix, correct the noise tokens, and replace the noise tokens with the categories predicted by the model. Assign weights to each sample, reduce the weights of noise samples, and use confidence or consistency scores as weights.
[0039] Retrain the model using the filtered dataset. Evaluate the classification accuracy and robustness of the model on the validation set or test set. Adjust the parameters (such as thresholds, number of clusters, etc.) of the noise filtering module according to the evaluation results. Train the model using the filtered dataset and evaluate its performance.
[0040] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application. Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concept. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the present application.
[0041] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. An online image classification method based on labeled distribution learning in an open environment, characterized in that: The method comprises the following steps: S1, build an image classification model based on labeled distribution learning, obtain an initial data set, perform feature increment on the initial data set, perform feature dimensionality reduction and feature selection on the accepted new features, enrich features and compress data to build a labeled distribution matrix to represent the relationship between images and categories, and obtain a new data set after feature dimensionality reduction; S2, for the new data set, update the model parameters online, mine the topological structure of the feature space, analyze the sample correlation through the sparse learning method, formulate the sample through the linear combination representation of the training samples, capture the mapping from the feature space to the label space, and migrate it to the tag space so that the tag space also satisfies a similar sparse structure; S3, design a dynamic category adjustment mechanism to automatically identify new categories and eliminate invalid categories; S4, introduces a noise filtering module to reduce the impact of labeling noise on the model; S5, retraining the model and evaluating it through the filtered data set, and adjusting the parameters of the noise filtering module according to the evaluation results, and finally obtaining the optimized model.
2. According to claim 1, the online image classification method based on labeled distribution learning in an open environment is characterized in that: The S1 specifically includes: S11, the initial data set is a small sample data set, record represents the q-dimensional sample space, represents the initial set of markers, the initial set of marker distributions, ,and Representation instance The initial data set includes labeled known data, unlabeled known data, and unlabeled unknown data, which are respectively denoted as ,in is a q-dimensional feature vector, expressed as ; S12, through the preset classifier The data is labeled and distributed and incorporated into the dataset In the example, the projection vector is recorded as p, and the feature vector of the sample can be obtained by the function Projected into a new eigenvector, expressed as a kernel function: ; S13, maps the marker distribution through a simple linear kernel function, ; Define the incremental kernel matrix of the feature space ,in ; Kernel matrix of label distribution ,in ; Here, the kernel function is replaced by the nuclear norm, which transforms it into a convex optimization problem: (1) in is the i-th singular value of matrix K, the nuclear norm It is a convex function, so we can find low-dimensional feature mapping by convex optimization; S14, Incremental HSIC in an open environment, , where tr(·) is the trace of the matrix, , e is a column vector of all 1s, F and G are the reproducing kernel Hilbert spaces mapped from X and d respectively. Substituting the above formula and removing the normalization term, we can get the optimal solution: (2) remember ,and , ; S15, after the incremental HSIC, the new data set after feature dimension reduction can be obtained ,in .
3. The online image classification method based on labeled distribution learning in an open environment according to claim 2, characterized in that: The S2 specifically includes: S21, find a suitable label distribution on the new dataset , so that it is The distance is the smallest, It is a sparse representation of sample points in the feature space, which is obtained by finding sparse representation in the feature matrix to explore the global structure of the feature space; S22, using the sparse representation method to model the relationship between a single example and other examples, thereby obtaining a sparse coefficient; S23, the quasi-Newton method is used to solve the optimal parameters.
4. The online image classification method based on labeled distribution learning in an open environment according to claim 3, characterized in that: The S3 specifically includes: S3.1, obtaining a sparse and well-structured adaptive graph, wherein the number of connected components in the adaptive graph is equal to the number of datasets / groups, S3.2, learning the adaptive graph, imposing structural constraints on the Laplacian graph, and introducing an adaptive structured graph for semi-supervised learning in an open environment; S3.3, combined with the maximum entropy model to optimize and update its parameters.
5. The online image classification method based on labeled distribution learning in an open environment according to claim 1, characterized in that: The S4 specifically includes: S41, construct the label distribution matrix and input the training data set ,in It is a sample. is a label; for each sample , calculate its association strength with all categories and generate a label distribution vector , where C is the number of categories; the label distribution vectors of all samples are combined into a label distribution matrix ; S42, calculating the statistical characteristics of each marker distribution vector, using a threshold or statistical test to identify outliers; clustering the marker distribution matrix, marking samples far from the cluster center as outliers; calculating the similarity between sample marker distribution vectors, and identifying marker distribution vectors that are significantly different from other samples as outliers; S43, using the confidence predicted by the model to evaluate the reliability of the marker, and identifying the marker with a confidence lower than a threshold as noise; using multiple models or the results of multiple trainings to perform consistency check, and identifying inconsistent markers as noise; analyzing abnormal patterns in the marker distribution matrix, and identifying markers that do not conform to the normal distribution as noise; S44, according to the model prediction results or the label distribution matrix, correct the noise labels and replace the noise labels with the categories predicted by the model; assign weights to each sample, reduce the weights of the noise samples, and use the confidence or consistency score as the weight.
6. The online image classification method based on labeled distribution learning in an open environment according to claim 1, characterized in that: The label distribution matrix is generated by the joint probability distribution of image features and category labels, and is used to represent the association strength between images and categories.
7. The online image classification method based on labeled distribution learning in an open environment according to claim 1, characterized in that: The model parameters are updated online using an incremental learning algorithm, which gradually optimizes the model parameters by minimizing the loss function.
8. The online image classification method based on labeled distribution learning in an open environment according to claim 1, characterized in that: The category dynamic adjustment mechanism automatically identifies new categories and removes invalid categories whose confidence is lower than a preset threshold by calculating the category confidence.
9. The online image classification method based on labeled distribution learning in an open environment according to claim 1, characterized in that: The noise filtering module identifies and filters the marker noise by analyzing the outliers in the marker distribution matrix.