Image classification method for guiding label relaxation and consistency supervision based on fuzzy label
By combining fuzzy C-means and true label supervision with a label relaxation mechanism, the problem of insufficient information integration between perspectives in multi-view image classification is solved, and the robustness and recognition ability of the model are improved. It is suitable for fields such as autonomous driving and remote sensing recognition.
Patent Information
- Application Number
- CN202510698802.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-19
AI Technical Summary
Existing multi-view image classification methods fail to effectively integrate the fuzzy structure and soft label information between perspectives, and have difficulty dealing with noisy samples and label uncertainty, resulting in limited recognition capabilities in complex scenes.
Fuzzy C-means is used to model the features of each viewpoint, obtain the fuzzy membership matrix, introduce real label supervision, and combine the label relaxation mechanism to construct the discriminative feature projection space, which is solved by a block-based alternating optimization strategy.
It improves the model's robustness to noise and uncertain data, enhances its ability to recognize complex scenes, reduces its dependence on high-quality labeled data, and improves training efficiency and computing resource utilization.
Smart Images

Figure CN120673130A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image classification, and in particular relates to an image classification method based on fuzzy label guided label relaxation and consistency supervision. Background Art
[0002] In recent years, multi-view data has been widely used in fields such as image recognition, behavior analysis, and scene classification. However, due to heterogeneity, feature redundancy, and fuzzy uncertainty between views, traditional single-view or simple fusion methods face performance bottlenecks in practical applications. In multi-view image classification tasks, in particular, the information provided by different viewpoints is complementary and locally different. How to effectively integrate features from each viewpoint, mitigate inconsistencies, and enhance the model's ability to discriminate against fuzzy samples has become a key research issue in this field.
[0003] Traditional multi-view classification methods rely heavily on hard-label supervision or explicit feature alignment mechanisms, failing to fully exploit the underlying fuzzy structure and soft label information between views. They often perform poorly with noisy samples or high data uncertainty. To address these issues, some research has begun to introduce fuzzy set theory to express the fuzzy membership of samples across multiple categories. However, most existing methods only use fuzzy modeling for unsupervised clustering or auxiliary feature selection, and have yet to systematically integrate fuzzy membership, ground-truth label supervision, and feature projection learning. This results in limited ability to collaboratively model fuzzy structure and semantic alignment.
[0004] The invention patent [Patent Number: CN202010440946.5] proposes a remote sensing image classification method using adaptive weighted multi-view metric learning. This method constructs a discriminative metric space from multiple viewpoints and introduces an adaptive adjustment mechanism for viewpoint weights, enabling the model to dynamically allocate weights across different viewpoints. Combined with kernel techniques, this method further improves classification performance in nonlinear spaces. Although this method has performed well in the field of remote sensing imagery, its primary reliance on the Mahalanobis distance metric and feature map optimization fails to effectively capture sample ambiguity, limiting its recognition capabilities when dealing with complex situations such as blurred class boundaries and uncertain labels. The invention patent [Patent Number: CN108021930B] proposes an adaptive multi-view image classification method and system. By combining multi-view label propagation with adaptive multi-graph weight learning, this method achieves fusion modeling of multi-view information within a unified semi-supervised learning framework. L2,1-norm sparse coding is used to construct weights for each viewpoint, effectively improving model robustness. However, this approach still primarily focuses on graph structure reconstruction and label propagation, lacking the ability to model the fuzzy attribution relationships between samples across viewpoints, making it difficult to address complex issues such as label uncertainty and multi-view conflicts. Furthermore, its weight optimization process lacks the guidance of real labels and does not incorporate a label relaxation mechanism to enhance semantic alignment. Summary of the Invention
[0005] Purpose of the invention: In response to the problems existing in the background technology, the present invention proposes a multi-view image classification method using fuzzy label guidance, label relaxation and consistency supervision from three directions: fuzzy label modeling, label relaxation guidance and consistency supervision. This method first uses fuzzy C-means to model the features of each view separately to obtain the fuzzy membership matrix, then introduces real label supervision to guide membership learning, and combines the label relaxation mechanism to construct a discriminative feature projection space, thereby improving the model performance while enhancing the robustness to noise and uncertainty data, and adapting to more complex real-world application scenarios.
[0006] Technical solution: An image classification method based on fuzzy label-guided label relaxation and consistency supervision of the present invention comprises the following steps:
[0007] Step 1: Collect multi-view image data of the target scene and divide it into training set and test set according to the proportion;
[0008] Step 2: Build an image classification model, take the training set as input, use fuzzy C-means to model the features of each viewpoint separately, and obtain the fuzzy membership matrix;
[0009] Step 3: Introduce real label supervision to guide membership learning in the learning process of fuzzy membership;
[0010] Step 4: Based on the fuzzy membership matrix and combined with the projection matrix learning of label relaxation, the discriminative mapping of multi-view data is achieved by minimizing the difference between the projected data and the relaxed labels;
[0011] Step 5: Use a block-based alternating optimization strategy to solve the image classification model, divide the overall complex optimization problem into decoupling sub-problems, and test the image classification model based on the test set to obtain the trained image classification model;
[0012] Step 6: Use the trained image classification model to classify the multi-view image data in the target scene.
[0013] Furthermore, step 2 is as follows: based on fuzzy set theory, a fuzzy C-means clustering strategy is introduced to describe the uncertainty of the sample category. The strategy allows samples to belong to multiple clusters at different degrees at the same time; the Euclidean distance is used to evaluate the correlation between samples and cluster centroids; if a sample is closer to a cluster centroid in the Euclidean distance, a higher h value can be obtained. ij The value is lower, and vice versa. The objective function is as follows:
[0014]
[0015] Among them jis the centroid of the jth cluster, h ij is the membership score between the i-th sample and the j-th cluster, H represents the membership matrix; to avoid extreme distribution of membership, a regularization constraint is added to the second term to ensure the sparse membership of each data point in different clusters, α is the regularization parameter;
[0016] For multi-view data Features from different perspectives X v (v=1,2,...,V) represents different specific attributes of the same sample. When analyzing data features from different perspectives, different results will be obtained. Further generalizing formula (1) to multiple perspectives will obtain multiple membership degrees and the following formula will be obtained:
[0017]
[0018] in It is the membership matrix calculated by the v-th view feature matrix.
[0019] In this way, the local structure of each view can be learned independently, effectively preserving the feature differences and fine-grained information under different views. At the same time, the fuzzy membership matrix H between multiple views v When subsequently combined with label information or feature projection process, it can also provide richer semantic clues.
[0020] Furthermore, step 3 is as follows: In the process of multi-view soft membership learning, it is often difficult to ensure consistency with the true category labels by relying solely on unsupervised clustering. Therefore, the real labels are introduced to guide the learning of each view membership, so that the model can better align with the real labels while retaining the flexibility of fuzzy clustering, thereby improving the discrimination and prediction accuracy. At the same time, adaptive view weights are introduced. It is used to represent the weight of the i-th sample in the v-th perspective. If a certain perspective is more advantageous for judging the sample, the model will tend to give it a higher value, otherwise it decreases. The learning function is:
[0021]
[0022] where Y i is the true label of the i-th sample, is the soft membership vector of the i-th sample at the v-th perspective, is the column vector consisting of the view weights corresponding to the i-th row membership matrix, is the view weight matrix corresponding to all data points.
[0023] By adding the guidance of real labels in the learning process of fuzzy membership, the model discrimination ability can be enhanced on the basis of unsupervised clustering. The adaptive adjustment fully utilizes the complementarity of data from different perspectives, making the membership distribution of samples among multiple clusters more flexible and better able to cope with noise and uncertainty.
[0024] Furthermore, step 4 is as follows: in the projection, the discriminative mapping of multi-view data is achieved by minimizing the difference between the projected data and the relaxed labels. Under this idea, the projection target is written as:
[0025]
[0026] Among them, P v Represents the projection matrix of the v-th perspective. In order to avoid overfitting in high-dimensional feature mapping, the projection matrix P v L2 regularization is applied to control the scale of the projection matrix coefficients, and γ is the regularization coefficient.
[0027] The projection method with additional "label relaxation" expressed in the above formula can make multi-view data closer to the true label after projection, while retaining the fuzziness and uncertainty characterized by the soft membership, providing a more flexible and accurate representation space for classification tasks.
[0028] Furthermore, step 5 is specifically as follows: the objective function of the optimization target algorithm is expressed as:
[0029]
[0030] Among them, β and λ are balance parameters and are greater than 0 to coordinate the influence of different parts of the objective function;
[0031] During optimization, an alternating optimization strategy is used to solve the optimization problem, optimizing the remaining variable by fixing the other variables.
[0032] Furthermore, the alternating optimization strategy is used to solve the optimization problem, and the remaining variables are fixed to optimize the remaining variable. The optimization process of each variable includes the following steps:
[0033] Step 5.1, fix H v ,P,W,Update
[0034] By fixing other variables, The optimization is converted into
[0035]
[0036] For each viewpoint, the optimization of each cluster centroid is independent, so The solution process is decomposed into V independent sub-problems:
[0037]
[0038] Since the objective function is convex, it is solved by setting its derivative to zero, so:
[0039]
[0040] Step 5.2, fix P, W, Update H v ;
[0041] By fixing other variables, H v The optimization derivation is:
[0042]
[0043] When only caring about the v-th perspective and the i-th sample When , the above formula is simplified to:
[0044]
[0045] in
[0046] make And expand formula (10) into The quadratic form of After removing the irrelevant constant terms, we can get
[0047]
[0048] Taking the derivative of the above formula and setting it to 0, we can get
[0049]
[0050] Arrangement available
[0051]
[0052] Step 5.3, Fixation H v ,P, update W;
[0053] By fixing other variables, W is optimized as follows:
[0054]
[0055] For each perspective, The optimization of are independent of each other. According to the constraints of formula (11), we have
[0056]
[0057] in
[0058] The Lagrangian function of formula (11) can be obtained as follows:
[0059]
[0060] where ψ is the Lagrange multiplier, which is obtained by using w i Calculate the derivative of equation (12) and set it to 0 to obtain w i The solution is
[0061]
[0062] Step 5.4, Fixation H v ,W, update P;
[0063] By fixing other variables, the optimization of P is derived as:
[0064]
[0065] This formula is equivalent to the following form
[0066]
[0067] At this time, take the partial derivative of the formula and set it to 0 to get
[0068]
[0069] Sort out the projection matrix P for each view
[0070] 2X ν (X ν P ν -(Y+H ν ))+2γP ν =0
[0071]
[0072] have to
[0073]
[0074] Algorithm 1 shows the algorithmic steps of the proposed FC-MVCC method.
[0075]
[0076]
[0077] The present invention further discloses a computer device, comprising a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method of the present invention.
[0078] The present invention further discloses a computer-readable storage medium having a computer program / instruction stored thereon, which implements the steps of the method of the present invention when the computer program / instruction is executed by a processor.
[0079] The present invention further discloses a computer program product, comprising a computer program / instruction, which implements the steps of the method of the present invention when executed by a processor.
[0080] Beneficial effects: Compared with the prior art, the present invention has the following significant advantages:
[0081] (i) The present invention introduces fuzzy C-means clustering to model the fuzziness and uncertainty in multi-view data, and represents the fuzzy attribution of samples to multiple categories through a soft membership matrix, thereby improving the classification system's ability to express complex sample structures.
[0082] (ii) An adaptive real label guidance mechanism is proposed to introduce real labels into the fuzzy membership learning process, and an automatic adjustment strategy for perspective weights is designed. This allows the model to more effectively align label information while taking into account the differences in features from different perspectives, thereby enhancing the discrimination ability.
[0083] (iii) A projection matrix learning method combined with a label relaxation mechanism is designed. By introducing soft membership into the projection space learning, a flexible semantic transition structure is constructed, which makes the projection features closer to the true label distribution and improves the generalization and stability of the model.
[0084] (iv) A block-based alternating optimization strategy is used to solve the model, dividing the overall complex optimization problem into decoupling sub-problems, improving training efficiency and significantly reducing computing resource consumption, making it suitable for large-scale data processing scenarios.
[0085] This method significantly improves the recognition ability of multi-view image classification systems under complex, high-dimensional, and multimodal data. It is particularly suitable for key areas such as autonomous driving, intelligent monitoring, and remote sensing identification, and enhances the intelligent system's perception and decision-making capabilities of the environment. By introducing fuzzy membership and label relaxation mechanisms, it effectively alleviates practical problems such as insufficient labels and strong data ambiguity, reduces dependence on high-quality labeled data, and thus saves a lot of labor costs. The model has strong robustness and stability, which can effectively reduce the risks brought by misclassification. BRIEF DESCRIPTION OF THE DRAWINGS
[0086] Figure 1 Flowchart of the present invention. DETAILED DESCRIPTION
[0087] The technical solution of the present invention is further described below.
[0088] Example
[0089] The present invention uses two real-world datasets, outdoor-scene and MSRC_v1, to verify the method of the present invention. The outdoor-scene dataset contains 8 outdoor scenes and 2688 color images, each of which is 256*256 pixels. For each image, four different feature vectors, 512-D GIST, 432-D color moment, 256-D HOG and 48-D LBP, are extracted to form a multi-view dataset; the MSRC_v1 dataset is a scene recognition dataset, containing 210 images from 7 categories of themes. In the experiment, a multi-view dataset was constructed by extracting five different features, CM, GIST, HOG, CENT and LBP, with the number of features for each view being 24, 576, 512, 256 and 254, respectively.
[0090] In the specific implementation, the dataset was split into 8:2, and no preprocessing was performed except for normalization. The parameters α, β, λ, and γ were all optimized using a 5-fold cross search + grid search within the range of {10-4, 10-2, 1, 102, 104}. The evaluation metric was set to Accuracy (ACC), and the average value was taken after 10 runs as the final experimental result. Figure 1 The experimental flow chart is shown below. In order to evaluate the prediction performance of the FC-MVCC algorithm for multi-view data, three advanced multi-view classification algorithms, IPMVSC, MvCCDA, and AWDR, and the traditional combination of non-negative matrix factorization (NMF) and support vector machine (SVM) for multi-view classification were selected as comparison algorithms. In addition, considering that the method proposed in this invention incorporates the idea of unsupervised clustering, the present invention also selected three multi-view classification algorithms using clustering algorithms, MLCK, DACK, and LACK, as comparison algorithms to test the effectiveness of consistency supervision.
[0091] Classification performance
[0092] In this section, the proposed method is compared with six multi-view algorithms and traditional classification algorithms. The results are shown in Table 1.
[0093] The method proposed in this paper achieved the best results on both the Outdoor-scene and MSRC_v1 datasets, with accuracy rates of 86.03%±1.59 and 97.07%±1.46, respectively, significantly outperforming all compared methods.
[0094] The Outdoor-scene dataset spans eight complex outdoor scenes and incorporates highly diverse viewpoint features such as GIST, color moment, HOG, and LBP. This poses a significant challenge to the feature fusion and robustness of multi-view methods. Compared to the superior IPMVSC (84.36% ± 0.37), our proposed method achieves an accuracy improvement of over 1.6% in this complex scene, demonstrating its fusion and generalization capabilities when dealing with high-dimensional, multimodal image features. Our proposed method also demonstrates strong performance on the MSRC_v1 dataset, which is characterized by small sample sizes, high semantic complexity, and strong viewpoint redundancy. Against this backdrop, our proposed method achieves an accuracy of 97.07% ± 1.46, a 2.91% improvement over the current state-of-the-art IPMVSC (94.16% ± 1.06). This demonstrates that the proposed fuzzy membership modeling and label relaxation mechanism effectively mitigate the impact of sample sparsity on multi-view fusion accuracy.
[0095] To further verify the statistical significance of this performance improvement, we performed a paired t-test to analyze the accuracy of each algorithm and our method. The results are shown in Table 2. Across all three datasets, our method demonstrated significant differences (p < 0.05) compared to most of the comparison methods. This demonstrates our method's excellent stability, adaptability, and discriminative power across a wide range of sample sizes, number of categories, and feature complexity.
[0096] Table 1 Comparative test results
[0097] Dataset Outdoor-scene MSRC_v1 IPMVSC 84.36±0.37 94.16±1.06 MvC 61.07±1.52 72.95±3.69 AWDR 82.16±1.31 88.34±4.05 NMF+SVM 84.13±1.74 92.38±2.78 MLCK 75.87±1.81 93.10±4.45 DACK 71.42±1.65 90.96±3.50 LACK 77.05±1.42 91.19±2.14 The method proposed by the present invention 86.03±1.59 97.07±1.46
[0098] Table 2 Statistical test results
[0099]
Claims
1. An image classification method based on fuzzy label guided label relaxation and consistency supervision, characterized in that: The steps include: Step 1: Collect multi-view image data of the target scene and divide it into training set and test set according to the proportion; Step 2: Build an image classification model, take the training set as input, use fuzzy C-means to model the features of each viewpoint separately, and obtain the fuzzy membership matrix; Step 3: Introduce real label supervision to guide membership learning in the learning process of fuzzy membership; Step 4: Based on the fuzzy membership matrix and combined with the projection matrix learning of label relaxation, the discriminative mapping of multi-view data is achieved by minimizing the difference between the projected data and the relaxed labels; Step 5: Use a block-based alternating optimization strategy to solve the image classification model, divide the overall complex optimization problem into decoupling sub-problems, and test the image classification model based on the test set to obtain the trained image classification model; Step 6: Use the trained image classification model to classify the multi-view image data in the target scene.
2. The image classification method based on fuzzy label guided label relaxation and consistency supervision according to claim 1, characterized in that: Step 2 is as follows: Based on fuzzy set theory, a fuzzy C-means clustering strategy is introduced to describe the uncertainty of the sample category. The strategy allows samples to belong to multiple clusters at different degrees at the same time; Euclidean distance is used to evaluate the correlation between samples and cluster centroids; If a sample is closer to a cluster centroid in Euclidean distance, a higher h can be obtained. ij The value is lower, and vice versa. The objective function is as follows: where x i represents the i-th sample, o j is the centroid of the jth cluster, h ij is the membership score between the i-th sample and the j-th cluster, and H represents the membership matrix; To avoid extreme distribution of membership, a regularization constraint is added in the second term to ensure the sparse membership of each data point in different clusters. α is the regularization parameter; For multi-view data Features from different perspectives X v (v=1,2,...,V) represents different specific attributes of the same sample. When analyzing data features from different perspectives, different results will be obtained. Further generalizing formula (1) to multiple perspectives will obtain multiple membership degrees and the following formula will be obtained: in It is the membership matrix calculated by the v-th view feature matrix.
3. The image classification method based on fuzzy label guided label relaxation and consistency supervision according to claim 1, characterized in that: Step 3 is as follows: introduce the real label to guide the learning of each view membership, so that the model can align the real label while retaining the flexibility of fuzzy clustering, and introduce adaptive view weights. It is used to represent the weight of the i-th sample in the v-th perspective. If a certain perspective is more advantageous for judging the sample, the model will tend to give it a higher value, otherwise it decreases. The learning function is: where Y i is the true label of the i-th sample, is the soft membership vector of the i-th sample at the v-th perspective, is the column vector consisting of the view weights corresponding to the i-th row membership matrix, is the view weight matrix corresponding to all data points.
4. The image classification method based on fuzzy label guided label relaxation and consistency supervision according to claim 3, characterized in that: Step 4 is as follows: In the projection, the discriminability of multi-view data is mapped by minimizing the difference between the projected data and the relaxed labels. The projection target is written as: Among them, P v Represents the projection matrix of the v-th perspective, in the projection matrix P v L2 regularization is applied to control the scale of the projection matrix coefficients, and γ is the regularization coefficient.
5. The image classification method based on fuzzy label guided label relaxation and consistency supervision according to claim 4, characterized in that: Step 5 is as follows: The objective function of the optimization target algorithm is expressed as: Among them, β and λ are equilibrium parameters and are greater than 0; An alternating optimization strategy is used to optimize the remaining variable by fixing the other variables.
6. The image classification method based on fuzzy label guided label relaxation and consistency supervision according to claim 5, characterized in that: The optimization using the alternating optimization strategy is specifically as follows: by fixing the remaining variables, the remaining variable is optimized. The optimization process of each variable includes the following steps: Step 5.1, fix H v ,P,W,Update By fixing other variables, The optimization is transformed into: For each viewpoint, the optimization of each cluster centroid is independent, so The solution process is decomposed into V independent sub-problems: Since the objective function is convex, it is solved by setting its derivative to zero, so: Step 5.2, fix P, W, Update H v ; By fixing other variables, H v The optimization derivation is: When only caring about the v-th perspective and the i-th sample When , the above formula is simplified to: in make And expand formula (10) into The quadratic form of After removing the irrelevant constant terms, we can get Taking the derivative of the above formula and setting it to 0, we can get Arrangement available Step 5.3, Fixation H v ,P, update W; By fixing other variables, W is optimized as follows: For each perspective, The optimization of are independent of each other. According to the constraints of formula (11), we have in The Lagrangian function of formula (11) can be obtained as follows: where ψ is the Lagrange multiplier, which is obtained by using w i Calculate the derivative of equation (12) and set it to 0 to obtain w i The solution is Step 5.4, Fixation H v ,W, update P; By fixing other variables, the optimization of P is derived as: This formula is equivalent to the following form At this time, take the partial derivative of the formula and set it to 0 to get Sort out the projection matrix P for each view 2X ν (X ν P ν -(Y+H ν ))+2γP ν =0 have to 7. A computer device comprising a memory, a processor, and a computer program stored in the memory, wherein: The processor executes the computer program to implement the steps of the method according to claim 1.
8. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to claim 1 are implemented.
9. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to claim 1 are implemented.
Citation Information
Patent Citations
An Adaptive Multi-View Image Classification Method and System
CN108021930B
A Remote Sensing Image Classification Method Based on Adaptive Weighted Multi-View Metric Learning
CN111680579B