Semi-supervised multi-mode classification method based on mode and strategy complementarity
Through the modal and strategic complementary learning framework, the modal interference and pseudo-label problems in semi-supervised multimodal classification are solved, and the classification accuracy is improved, which is suitable for multimodal classification tasks in the image-text field.
Patent Information
- Application Number
- CN202510403368.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
AI Technical Summary
In the existing semi-supervised multimodal classification methods, information imbalance in image and text modality leads to modal interference problems, pseudo-label prediction errors, limiting classification performance, and the fractional fusion strategy retains the correct number of pseudo-labels, while the feature splicing strategy has a low accuracy rate.
The modal complementary learning module and the strategy complementary learning module are adopted to automatically learn modal reliability weights through modal reliability generator and pseudo-label consistency guidance, combining weighted fraction fusion and feature splicing, and using modal and strategy complementarity to improve the trade-off between the quality and quantity of pseudo-labels.
It effectively improves the accuracy of semi-supervised multimodal classification, solves the modal interference problem, and expands the application value of a variety of fusion methods and pseudo-label methods in practical application scenarios.
Smart Images

Figure CN120336954A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computers, and more particularly, to a semi-supervised multi-modal classification method based on modality and policy complementarity. Background Art
[0002] Multi-modal classification has attracted great interest in the academic community. Traditional supervised models require a sufficient amount of labeled data and outperform single-modal classifiers. However, in the real world, the cost of collecting and labeling large amounts of multi-modal data is high. A practical solution is semi-supervised learning, which can effectively utilize easily accessible unlabeled data. In the present invention, a task called semi-supervised multi-modal classification (SSMC) is explored, which extends semi-supervised learning to multi-modal classification scenarios involving image and text modalities.
[0003] Common methods for semi-supervised classification are to use regularization information. For example, UDA and MixText strengthen consistency training by augmenting data. Another idea is to use pseudo-labels during training to increase the amount of labeled data. For example, a series of works starting from MixMatch label unlabeled data and select reliable data at a fixed or variable threshold. There are also some methods that utilize unlabeled data by combining pseudo-labels with prototype learning, or construct a dictionary according to attention. Some semi-supervised ideas have been applied to the audio and video fields. Some methods focus on crisis analysis tasks.
[0004] Multi-modal classification has been studied in the image-text field. Mainstream methods include late fusion and early fusion. Late fusion uses encoders of different modalities and fuses the extracted features. For example, UniS-MMC designs a splicing and comparison learning mechanism; others design different modalities to be embedded in a cross-attention mechanism. Early fusion focuses on the association between different modal network layers. For example, MI2P designs fine-grained plug-and-play modules before encoding features. Prior work has also studied the importance of different modalities based on prior knowledge of modal strength or validation set accuracy.
[0005] However, in the field of graphics and text, there are few studies on multimodal classification tasks under semi-supervised conditions. In order to study the problem, the present invention first combines the existing semi-supervised methods and modal fusion methods to design a baseline semi-supervised multimodal model. The present invention selects the most common semi-supervised learning method-pseudo-labeling method as the backbone. Specifically, the famous FixMatch and FreeMatch are used as benchmark semi-supervised learning methods. In terms of modal fusion, the classic label-level aggregation method score fusion strategy (Score Fusion) and the feature-level fusion method feature concatenation strategy (Feature Concat) are adopted. The former is used to fuse the prediction results of different modalities, and the latter first concatenates the features of different modalities and then classifies them.
[0006] Based on the study of semi-supervised multimodal classification and baseline models, the present invention discovered the following two phenomena. First, both image and text modes are helpful for classification, and they complement each other in the semi-supervised multimodal classification task. The information provided by these two modes is unbalanced, which may cause modal interference problems when applying average score fusion. In this case, when the prediction of one of the modalities is incorrect, even if the prediction of the other modality is correct, the fused prediction may be wrong. Figure 1 As shown in Figure 2, this problem often occurs when the classification results of different modalities are inconsistent, and the average score can easily lead to incorrect fusion results. This phenomenon may lead to incorrect pseudo-label predictions and limit the performance of semi-supervised multimodal classification.
[0007] Secondly, pseudo labels of different strategies show complementarity in semi-supervised multimodal classification. Figure 1 As shown in Figure 3, when using thresholds to filter pseudo-labels, the score fusion strategy can retain the correct pseudo-labels with high accuracy, but the number of selected instances is significantly less than the feature splicing strategy. The opposite is true for the feature splicing strategy, which can retain more pseudo-labeled data but has lower accuracy. Since both the quality and quantity of pseudo-label examples are crucial in semi-supervised learning, it is necessary to consider the complementarity between the two strategies when building semi-supervised multimodal classification models. Summary of the invention
[0008] The purpose of the embodiments of the present disclosure is to provide a semi-supervised multimodal classification method based on the complementarity of modalities and strategies, to solve the problem that the existing pseudo-label prediction is wrong and limits the performance of semi-supervised multimodal classification, and when using thresholds to filter pseudo-labels, the score fusion strategy can retain the correct pseudo-labels with a high accuracy, but the number of selected instances is significantly less than the feature splicing strategy. The situation of the feature splicing strategy is the opposite, it can retain more pseudo-labeled data, but the accuracy is low.
[0009] In one general aspect, a semi-supervised multi-modal classification method based on modality and policy complementarity is provided, including: a modality complementary learning module and a policy complementary learning module; during the classification inference process, this method needs to simultaneously receive the image and text in the input image-text pair data and hand them over to the encoders of their respective modalities, where:
[0010] For the input image-text pair data, first, through the modality complementary learning module, apply the encoders of text and image respectively, and use an MLP model with two softmax functions to obtain the probability p, and perform single-modal prediction on the text or image data respectively; then, based on the extracted features or single-modal prediction results, obtain the results of the weighted score fusion method and the results of feature concatenation;
[0011] And use the results of the weighted score fusion and the feature concatenation results to input into the policy complementary learning module to form a baseline model for semi-supervised multi-modal classification;
[0012] The supervised loss L of the baseline model for semi-supervised multi-modal classification sup is expressed as:
[0013]
[0014] where B x is the batch size of the sampled labeled data;
[0015] In the output part, use the label-level fusion method and the feature-level fusion method to obtain two different prediction results, take the average of their prediction probabilities to obtain the final class probability, and select the class with the highest probability as the classification result, where:
[0016] For each unlabeled data u i assign a pseudo-label to it, and prepare a weakly augmented version a(u i ) and a strongly augmented version A(u i ), and determine the pseudo-label of u i through , and select the pseudo-label with high confidence to calculate the loss L i : u :
[0017]
[0018] where B u is the batch size of unlabeled data, and τ is a fixed or variable threshold for filtering low-quality pseudo-labels;
[0019] The pseudo-label method uses L sup +λL u or adds additional regularization as the optimized loss.
[0020] The specific method of the weighted score fusion is as follows: introduce the MLP model M w , and use the softmax function to obtain the reliability distribution w(z|t,v) of different modalities from the concatenated features [f(t), g(v)]. The specific method is as follows:
[0021]
[0022] where is the parameter corresponding to the modality z';
[0023] For each image-text pair (t,v), adopt the weighted fusion function:
[0024] p S (y|t,v) = w(z = 0|t,v)p T (y|t) + w(z = 1|t,v)p V (y|v)
[0025] This function is used for the score fusion branch.
[0026] To address the different training convergence speeds of strong and weak modalities, design a label consistency guidance method to calibrate the weight learning process using limited labels. Use pseudo-labels and labeled data in the same batch to form the dataset D sub in a batch to train the modality reliability generator;
[0027] The specific process is as follows: First, obtain the prediction results through strong augmentation of the unlabeled input data:
[0028]
[0029] where K is the set of categories, k is a certain category in the category set; A(t) is the strongly augmented version of the text, A(v) is the strongly augmented version of the image, and p T (k|A(t)) is the category prediction probability obtained from the strongly augmented version of the text, and p V (k|A(v)) is the category prediction probability obtained from the strongly augmented version of the image; r t is the predicted category result obtained from the text modality, and r v is the category prediction result obtained from the image modality.
[0030] Then, generate the supervision signal according to the consistency of the pseudo-labels generated by the strongly augmented version and the weakly augmented version of the input:
[0031]
[0032] where G(z|t,v) is the pseudo-label reliability distribution probability, is a pseudo-label, is the probability predicted as the category according to the text input , is the probability predicted as the category according to the image input ; r t and r v is the category prediction result obtained from the above process.
[0033] Correspondingly, there is G(z = 1|t, v) = 1 - G(z = 0|t, v);
[0034] When r t or r v in one of the modalities is predicted to be consistent with the pseudo-label , the modality with the consistent prediction is regarded as a reliable modality, and thus a probability score of 1 is assigned to this modality;
[0035] When the prediction results r t and r v of both modalities are inconsistent with the pseudo-label , the prediction results of both modalities are more likely to be unreliable, and the contributions of both modalities are evenly divided;
[0036] When the modality predictions r t and r v are both consistent with the pseudo-label , the modality prediction with a higher confidence is regarded as a more reliable prediction, and the probability score is calculated by performing a softmax operation based on their corresponding maximum prediction values;
[0037] The KL divergence is used to make the generated reliability distribution w(z|t, v) close to the guidance value G(z|t, v), and this part of the loss L lcg is expressed as follows:
[0038]
[0039] where |D sub | is the size of the subset of the joint set selected in the batch.
[0040] To solve the problem that the reliability weight is small and the gradient generated by the unreliable modality is not sufficient to train the corresponding branch, a modality reliability guidance module is set in the modality complementary learning module, and the same data D sub as the label consistency guidance is adopted. The data subset is divided into two parts, and D t2i ∈ D sub is introduced to represent the data set where the text modality is more reliable, and Di 2i ∈ D sub is introduced to represent the data set where the image modality is more reliable; for each data and label (t, v, y) in D sub , if rt is consistent with y, while r v is inconsistent with y, or both are consistent with y, but the maximum predicted value of p T is higher than p V and add this data to D t2i ;
[0041] Use the KL divergence to complete the training process and introduce L mrg to represent the loss of the modality reliability guidance module:
[0042]
[0043] where |D t2i | and |D i2t | are the sizes of the two sets.
[0044] The strategy complementary learning module introduces a consistency pseudo-label selection method to construct a dataset for guiding information; first, aggregate the labeled data and unlabeled data with high confidence. For the unlabeled data, retain the data where the pseudo-labels predicted by the two strategies are consistent and the maximum probability of one of them is higher than the threshold η, and use U sub to represent the subset of unlabeled data in a batch B as:
[0045]
[0046] where, and are the pseudo-labels of weighted score fusion and feature splicing;
[0047] Then, use X sub ={(x i ,y i )|(x i ,y i )∈B} to represent all the labeled data in the same batch, and the union set is D sub =U sub ∪X sub ;
[0048] Design a strategy complementary guidance module to apply the strategy complementarity of weighted score fusion and feature splicing. Use D s2c ∈D sub to represent the data learned by feature splicing from score fusion, use D c2s ∈D sub to represent the opposite data, and use strong augmentation to collect and where K is the set of classes, k is a certain class in the class set; A(t) is the strongly augmented version of the text, A(v) is the strongly augmented version of the image; p S(k|(A(t),A(v)) is the predicted probability obtained by the fractional fusion method, p C (k|(A(t),A(v)) is the predicted probability obtained by the feature splicing method; r s is the predicted class result obtained by the fractional fusion method, r C is the predicted class result obtained by the feature splicing method; as the predicted results of unlabeled data for weighted fractional fusion and feature splicing, and r of labeled data is obtained using unenhanced data s and r C ; D s2c and D c2s The rule of is expressed as: for each data and label (t, v, y) in D sub , if r s is consistent with y, but r C is inconsistent with y, or both are consistent with y, but the maximum predicted value of p S is higher than p C , add these data to D s2c , and promote p C to learn from p S ;
[0049] After that, use the KL divergence to force the weaker-performing strategy to learn from the better-performing strategy, so as to obtain L scg :
[0050]
[0051] where, |D s2c | and |D c2s | are the sizes of the two sets.
[0052] The technical effects to be achieved by the embodiments of the present invention are as follows:
[0053] There are two complementarities widely existing in semi-supervised multi-modal tasks, namely modal complementarity and strategy complementarity. The present invention proposes an application framework in semi-supervised multi-modal tasks based on these two complementarities, and can be used in combination with many existing semi-supervised pseudo-label methods, effectively improving the classification accuracy. In actual application scenarios, since it is very difficult to collect accurately labeled data under multi-modal conditions, therefore, this solution can be effectively extended to actual application scenarios, and can be combined with a variety of different fusion methods and semi-supervised pseudo-label methods, having quite strong application value.
[0054] In summary, the work of the present invention has the following contributions: (1) The present invention introduces the semi-supervised multi-modal classification problem and reveals the complementarity that helps to improve semi-supervised multi-modal classification. (2) The present invention proposes a modal complementary learning method to estimate modal reliability and reduce the modal interference problem. (3) The present invention proposes a strategy complementary learning method to balance the quality and quantity of pseudo-labels. (4) The present invention creates three benchmarks for semi-supervised multi-modal classification to evaluate the framework of the present invention. Experimental analysis proves the effectiveness of the method of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] The above and other objects and features of the present disclosure will become more apparent from the following description with reference to the accompanying drawings.
[0056] Figure 1 is a schematic diagram showing the problems existing in semi-supervised multi-modal classification in the prior art according to the present disclosure;
[0057] Figure 2 is a schematic diagram showing a semi-supervised multi-modal classification method based on motif and strategy complementarity according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0058] The following detailed description is provided to assist the reader in obtaining a comprehensive understanding of the methods, apparatuses, and / or systems described herein. However, various changes, modifications, and equivalents of the methods, apparatuses, and / or systems described herein will be apparent after understanding the disclosure of the present application. For example, the order of operations described herein is merely exemplary and is not limited to those set forth herein, but may be changed as will be apparent after understanding the disclosure of the present application, except for operations that must occur in a specific order. In addition, descriptions of features known in the art may be omitted for greater clarity and conciseness.
[0059] The features described herein may be implemented in different forms and should not be construed as limited to the examples described herein. Instead, the examples described herein are provided only to illustrate some of the many possible ways of implementing the methods, apparatuses, and / or systems described herein, which will be apparent after understanding the disclosure of the present application.
[0060] As used herein, the term "and / or" includes any one of the associated listed items and any combination of any two or more thereof.
[0061] Although terms such as "first", "second", and "third" may be used herein to describe various components, elements, regions, layers, or sections, these components, elements, regions, layers, or sections should not be limited by these terms. Instead, these terms are only used to distinguish one component, element, region, layer, or section from another. Thus, a first component, first element, first region, first layer, or first section described in the examples herein may also be referred to as a second component, second element, second region, second layer, or second section without departing from the teachings of the examples.
[0062] In the specification, when an element (such as a layer, region, or substrate) is described as "on" another element, "connected to" or "coupled to" another element, the element may be directly "on" the other element, directly "connected to" or "coupled to" the other element, or there may be one or more other elements therebetween. In contrast, when an element is described as "directly on" another element, "directly connected to" or "directly coupled to" another element, there may be no other elements therebetween.
[0063] The terms used herein are only for describing various examples and are not intended to limit the disclosure. Unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. The terms "comprising", "including", and "having" specify the presence of the recited features, quantities, operations, components, elements, and / or combinations thereof, but do not preclude the presence or addition of one or more other features, quantities, operations, components, elements, and / or combinations thereof.
[0064] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains after understanding this disclosure. Unless explicitly defined herein, terms (such as those defined in a general dictionary) should be construed to have a meaning consistent with their meaning in the context of the relevant art and this disclosure, and should not be interpreted in an idealized or overly formal manner.
[0065] In addition, in the description of the examples, when a detailed description of a related structure or function that is considered to be well-known would cause an ambiguous interpretation of the disclosure, such a detailed description will be omitted.
[0066] Figure 1 is a schematic diagram showing an architecture of a semi-supervised multi-modal classification method based on modal and strategy complementarity according to an embodiment of the present disclosure.
[0067] Since the performance of semi - supervised multi - modal classification tasks can be improved by leveraging the complementarity between different modalities and strategies, the present invention proposes a framework for learning from modality and strategy complementarity (MSC). Specifically, the modality complementarity learning module of the present invention aims to assign a reliability weight to each modality, as Figure 1 shown. Different from previous works that regarded modality weights as hyperparameters or calculated weights based on the validation set, the present invention learns modality reliability weights from different individual instances and calibrates them according to the consistency of predictions. The present invention further utilizes the KL divergence to force the modalities with lower reliability to learn from the modalities with higher reliability at the instance level, thereby reducing the error accumulation during the training process. By this method, the present invention successfully solves the modality interference problem, thus improving the effect of semi - supervised multi - modal classification.
[0068] To learn from strategy complementarity, the present invention proposes a method for selecting pseudo - labeled data, which utilizes the advantages of two strategies: score fusion and feature concatenation. Specifically, this method uses the combination of the data selected by these two strategies to train the modality and strategy complementarity learning framework. The selected data is regarded as a guiding signal in the bidirectional KL divergence loss, allowing the weaker strategy branch to learn from the stronger strategy branch. This learning process utilizes the complementarity between strategies and makes a trade - off between the quality and quantity of pseudo - labeled data, thus contributing to improving the performance of semi - supervised multi - modal classification tasks.
[0069] Considering that the research on semi - supervised multi - modal classification tasks in the current image - text field is not sufficient and there is no widely recognized semi - supervised multi - modal classification test dataset, in order to effectively measure the effectiveness of the modality and strategy complementarity learning framework studied in the present invention in semi - supervised multi - modal classification tasks, the present invention creates semi - supervised multi - modal benchmarks with different settings based on three commonly used datasets in the image - text field. Specifically, the present invention retains the categories with a larger amount of data, randomly selects 20, 50, 100 labeled data for each category, and retains 500 - 2000 unlabeled data for each category according to the size of the specific dataset to simulate a relatively common semi - supervised environment. By comparing with the supervised and unsupervised baseline models constructed by the present invention in terms of classification accuracy and F1 - score metrics, the present invention
[0070] Preliminary work:
[0071] Problem definition:
[0072] Semi-supervised multi-modal classification aims to classify image-text based on a training set consisting of a small amount of labeled data and a large amount of unlabeled data. Formally, let \(Z = \{0, 1\}\) be the modal set, and \(K\) represent the class set. In the present invention, \(Z\) includes text and image modalities. Let \(D\) be the training set, which contains a labeled set \(X\) and an unlabeled set \(U\). The labeled set is represented as where and refer to the text and image of the \(i\)-th data pair, and \(y\) i is its corresponding label. The unlabeled set is represented as where and represent the \(i\)-th unlabeled image-text pair. For simplicity, each labeled and unlabeled image-text pair is also represented as \(x\) i and \(u\) i . The goal of semi-supervised multi-modal classification is to predict the class of a given data pair through the finite labeled data set \(D\).
[0073] Multi-modal classification baseline model:
[0074] Followed the settings of previous work, that is, using the BERT model to encode text and the ViT model to encode images. Specifically, let \(f(t)\) and \(g(v)\) represent the encoders of text \(t\) and image \(v\) respectively, and an MLP model \(M\) with two softmax functions T and \(M\) V are used to obtain the probability \(p\):
[0075]
[0076] where, and are the parameters corresponding to class \(k\). Using the models in the above equations, unimodal predictions can be made for text or image data respectively.
[0077] Based on the extracted features or unimodal prediction results, score fusion and feature concatenation can be used to perform multi-modal classification. Score fusion is to average the unimodal predictions:
[0078] \(p\) S (y|t,v) = (p T (y|t) + p V (y|v)) / 2
[0079] Feature concatenation first concatenates the features extracted from the text and image modalities, and then performs softmax classification on the concatenated features. Specifically, an MLP model \(M\) C and a softmax function are used to obtain the prediction result:
[0080]
[0081] Among them, is a parameter corresponding to class k. Score fusion and feature concatenation are the most commonly used modality fusion strategies for multi-modal classification, and both of them will be used to construct a baseline model for semi-supervised multi-modal classification.
[0082] Semi-supervised classification baseline model:
[0083] Select the commonly used semi-supervised classification models FixMatch and FreeMatch as benchmarks. Both of these models contain losses for labeled data x i and unlabeled data u i respectively, where the supervised loss L sup can be expressed as:
[0084]
[0085] where B x is the batch size of the sampled labeled data.
[0086] For each unlabeled data u i , a pseudo-label will be assigned to it and used as additional training data. A weakly augmented version a(u i ) and a strongly augmented version A(u i ) are prepared for u i . The prediction of a(u i ) is used to select high-quality pseudo-labeled data, and the cross-entropy loss with A(u i ) is calculated. Specifically, through to determine the pseudo-label of u i , and high-confidence pseudo-labels are selected to calculate the loss L u :
[0087]
[0088] where B u is the batch size of the unlabeled data, and τ is a fixed or variable threshold used to filter low-quality pseudo-labels.
[0089] The pseudo-label method uses L sup +λL u or adds additional regularization as the optimized loss. Using L S and L C to represent the losses when p is specified as p S and p C respectively. The baseline multi-modal classification model is combined with known semi-supervised methods to create a baseline semi-supervised multi-modal classification model.
[0090] By analyzing the baseline model, two complementary properties were discovered, which can be utilized to improve the semi-supervised multi-modal classification task. First, the predicted values p T and p V from different modalities are complementary. Some predictions are dominated by one of the modalities, and the existing average score fusion may encounter the modality interference problem, resulting in incorrect fused scores in the predictions (as shown in the upper half of Figure 1 ). To address this issue, a method of learning from modality complementarity was proposed, that is, automatically learning reliability weights to weigh the prediction results from different modalities and correct the fused scores.
[0091] Second, empirical studies have shown that when selecting pseudo-labeled data, the score fusion strategy and the feature concatenation strategy are complementary in terms of quality and quantity. Specifically, as shown in the lower half of Figure 1 the score fusion p S tends to select fewer but higher-quality pseudo-labeled data. In contrast, the feature concatenation p C is prone to selecting more but lower-quality pseudo-labeled data. Since the quality and quantity of pseudo-labeled data are both important for effective semi-supervised learning, there is an incentive to learn from the strategy complementarity, so that a trade-off can be made between the quality and quantity of the selected pseudo-labeled data, enabling different strategies to complement each other's strengths and weaknesses.
[0092] Empirical studies on the baseline model have shown that there are modality complementarity and strategy complementarity in semi-supervised multi-modal classification, and appropriately leveraging these two complementary properties can improve performance. Based on this, a modality and strategy complementarity learning framework was designed (the structure is as shown in Figure 2 ), and how to utilize these two complementary properties was introduced.
[0093] Learning from modality complementarity:
[0094] As analyzed previously, to address the modality interference problem in semi-supervised multi-modal classification mentioned above, a modality reliability generator was designed to estimate the reliability of different modality predictions and use the stronger modality to guide the weaker modality.
[0095] The solution to the modality interference problem is to assign weights representing reliability to each modality and convert the average score fusion into a weighted score fusion. To obtain the modality reliability weights, a modality reliability generator was proposed to generate the reliability scores for each modality. Specifically, an MLP model M w was introduced, and then using the softmax function, the reliability distribution w(z|t,v) of different modalities was obtained from the concatenated features [f(t),g(v)] as follows:
[0096]
[0097] Among them, is a parameter corresponding to the modality z'. Use w(z = 0|t, v) and w(z = 1|t, v) to represent the reliability of the text modality and the image modality. For each image-text pair (t, v), modify the average score fusion function to a weighted fusion function:
[0098] p S (y|t, v) = w(z = 0|t, v)p T (y|t) + w(z = 1|t, v)p V (y|v)
[0099] This weighted fusion method can alleviate the modality interference problem by assigning smaller weights to the predictions of unreliable modalities, thereby improving the performance of semi-supervised multi-modal classification. In the model framework, the above formula is used for the score fusion branch.
[0100] As mentioned in previous work, the training convergence speeds of strong modalities and weak modalities are different, which will cause the modality weights to bias towards strong modalities, thus destroying the weighted fusion. To solve this problem, a Label Consistency Guidance (LCG) method is designed to calibrate the weight learning process using limited labels. Use the pseudo-labels collected in policy complementary learning and the labeled data in the same batch to form a dataset D sub in a batch to train the modality reliability generator. The specific rules will be introduced in the "Learning from Policy Complementary" section. First, obtain the prediction results by strongly augmenting the unlabeled input data, as follows:
[0101]
[0102] Among them, K is the set of classes, k is a certain class in the class set; A(t) is the strongly augmented version of the text, A(v) is the strongly augmented version of the image, p T (k|A(t)) is the class prediction probability obtained from the strongly augmented version of the text, p V (k|A(v)) is the class prediction probability obtained from the strongly augmented version of the image; r t is the predicted class result obtained by the text modality, r v is the class prediction result obtained by the image modality.
[0103] Then, generate a supervision signal according to the consistency of the pseudo-labels generated by the strongly augmented version and the weakly augmented version of the input, which can be expressed as follows:
[0104]
[0105] Among them, G(z|t,v) is the probability distribution of the reliability of the pseudo-label, is the pseudo-label, is the probability predicted as class according to the text input, is the probability predicted as class according to the image input; r t and r v are the class prediction results obtained from the above process.
[0106] Correspondingly, we have G(z = 1|t,v) = 1 - G(z = 0|t,v). The intuitive explanation of G is as follows:
[0107] When the prediction of one of the modalities in r t or r v is consistent with the pseudo-label , the modality with the consistent prediction is regarded as a reliable modality, and thus a probability score of 1 is assigned to this modality, and vice versa.
[0108] When the prediction results of both modalities r t and r v are inconsistent with the pseudo-label , the prediction results of both modalities are more likely to be unreliable. Therefore, let the contributions of both modalities be evenly divided.
[0109] When the modality predictions r t and r v are both consistent with the pseudo-label , the modality prediction with a higher confidence is regarded as a more reliable prediction. Therefore, the probability scores are calculated by performing a softmax operation based on their corresponding maximum prediction values.
[0110] The KL divergence is used to make the generated reliability distribution w(z|t,v) close to the guidance value G(z|t,v). This part of the loss L lcg is expressed as follows:
[0111]
[0112] where |D sub | is the size of the subset of the joint set selected in the batch. It should be noted that the ground truth label y is used to replace and the prediction of the labeled data is obtained using the non-augmented data (t,v) ∈ X.
[0113] Since the reliability weights are small, the gradients generated by unreliable modalities may not be sufficient to train the corresponding branches. Therefore, unreliable modalities are allowed to learn from more reliable modalities according to the instance-level reliability, and the Modality Reliability Guidance (MRG) module is proposed.
[0114] Adopt the same data D as the label consistency guidance sub . The data subset is divided into two parts. Formally, introduce D t2i ∈D sub to represent the dataset with more reliable text modality, and D i2t ∈D sub to represent the dataset with more reliable image modality. The prediction results of these two groups of data for the text modality r t and the results of the image modality r v are explained as follows:
[0115] For each data and label (t, v, y) in D sub , if r t is consistent with y, while r v is inconsistent with y, or both are consistent with y, but the maximum predicted value of p T is higher than that of p V , add these data to D t2i , which forces p V to learn from p T on these data, and vice versa.
[0116] Here, the KL divergence is used to complete the training process. Introduce L mrg to represent the loss of the modality reliability guidance module as follows:
[0117]
[0118] where |D t2i | and |D i2t | are the sizes of the two sets. Through the above design, appropriate reliability weights are assigned to different modalities, solving the modality interference problem in semi-supervised multi-modal classification.
[0119] Learning from policy complementarity:
[0120] Different fusion strategies have different performances in the pseudo-label selection of semi-supervised multi-modal classification, which shows policy complementarity. To make full use of this, a consistency pseudo-label selection method is introduced to construct a dataset for guiding information.
[0121] To make full use of the unlabeled data, it is recommended to aggregate the labeled data and unlabeled data with high confidence according to the consistency of the two strategies. For the unlabeled data, retain the data where the pseudo-labels predicted by the two strategies are consistent and the maximum probability of one of them is higher than the threshold η. Specifically, use U sub to represent the unlabeled data subset of a batch B as:
[0122]
[0123] Among them, and are the pseudo-labels for weighted score fusion and feature concatenation. This choice can be regarded as a voting ensemble and can achieve a high accuracy. Then, use X sub ={(x i ,y i )|(x i ,y i )∈B} to represent all the labeled data in the same batch, and the union set is D sub =U sub ∪X sub . This set is used in the previous section to form the guidance loss and will be used in the policy complementary guidance module.
[0124] To utilize the policy complementarity of weighted score fusion and feature concatenation introduced in the previous sections, a Policy Complementary Guidance (SCG) module is proposed to promote the two policies to learn from each other at the instance level and improve the semi-supervised multi-modal classification performance. Specifically, use D s2c ∈D sub to represent the data that feature concatenation learns from score fusion, and use D c2s ∈D sub to represent the opposite data. Similarly, use strong augmentations to collect as well as as the predicted results of the unlabeled data for the two policies, where K is the set of classes, k is a certain class in the set of classes; A(t) is the strongly augmented version of the text, A(v) is the strongly augmented version of the image; p S (k|(A(t),A(v)) is the predicted probability obtained by the score fusion method, p C (k|(A(t),A(v)) is the predicted probability obtained by the feature concatenation method; r s is the predicted class result obtained by the score fusion method, r C is the predicted class result obtained by the feature concatenation method. And use the non-augmented data to obtain r s and r C of the labeled data. The rules for D s2c and D c2s are expressed as:
[0125] For each data and label (t, v, y) in D sub , if r s is consistent with y, but r C is not consistent with y, or both are consistent with y, but the maximum predicted value of p S is higher than p C , add these data to D s2c to push p C away from pS Learn from it, and vice versa.
[0126] Use the KL divergence to force the weaker-performing strategy to learn from the better-performing strategy, thus obtaining L scg :
[0127]
[0128] Among them, |D s2c | and |D c2s | are the sizes of two sets. Through the strategy complementary guidance module, score fusion can obtain more pseudo-label data while maintaining the original quality, while feature concatenation can obtain higher-quality pseudo-labels while maintaining the original quantity. Strategy complementarity makes semi-supervised multi-modal classification make a trade-off between the quantity and quality of pseudo-labels. It should be noted that after applying the modal reliability weights, the performance of weighted score fusion in selecting pseudo-label data is still similar to that of average score fusion, so this guidance rule is still applicable.
[0129] Training objective:
[0130] Use the above three loss functions, combined with the basic pseudo-label semi-supervised methods FixMatch and FreeMatch, to train the framework as follows:
[0131] L = L S + L C + β1L lcg + β2L mrg + β3L scg
[0132] Among them, β1, β2, β3 are weight coefficients used to balance the influence of different losses. During the evaluation process, take the average of the framework prediction probabilities p S and p C as the final integration result.
[0133] Datasets and experimental settings:
[0134] Select three datasets to construct the benchmark test datasets for semi-supervised multi-modal classification: N24News, UPMC-Food101, CrisisMMD. Concatenate the titles, abstracts, and captions of N24News to obtain the full text. Sample data pairs with the same labels on two modalities and evaluate only on Task 1 of CrisisMMD. Divide the training set into 20, 50, 100 labeled data pairs. The dataset statistics are shown in Table 1.
[0135] Table 1 Dataset statistics
[0136]
[0137] The accuracy and F1 are used to evaluate the performance of semi - supervised multi - modal classification. SF and FC are used as abbreviations for score fusion and feature concatenation respectively.
[0138] Baseline models:
[0139] The BERT model for text and the ViT model for images are selected as single - modal supervised baseline models. For the multi - modal supervised model, score fusion and feature concatenation are selected as multi - modal baselines. To confirm the generality of the framework in pseudo - labeling methods, two semi - supervised methods, FixMatch and FreeMatch, are selected and combined with the above - mentioned methods as semi - supervised baseline models.
[0140] Implementation:
[0141] BERT and ViT are used to encode text and images. The framework is combined with FixMatch and FreeMatch. The batch size B of labeled data x is set to 4, and the batch size B of unlabeled data u is set to 32. We set the learning rate to 5e - 5. RandAugment is used to obtain augmented images, and augmented text is obtained through swapping and synonym strategies. Sentence - bert is used to calculate the similarity between the augmented text and the original text, and the text with higher similarity is selected as weakly augmented text, and the other text is selected as strongly augmented text. AdamW is selected as the optimizer. η is set to 0.95. The number of training epochs is set to 20. For N24News and UPMC - Food101, β1, β2, β3 are set to 1; for CrisisMMD, they are set to 0.6.
[0142] Main results:
[0143] The semi - supervised multi - modal classification results are shown in Table 2. The above four supervised methods are trained only using labeled data. It is observed that the accuracy and F1 of the modality and strategy complementary learning framework are generally better than those of the supervised and semi - supervised baselines. Especially on UPMC - Food101, the framework is about 4% - 5% better than the best baseline. On CrisisMMD and N24News, it is also 0.5% - 3% higher than the baseline. This shows the effectiveness of the modality and strategy complementary learning framework combined with semi - supervised methods. Another conclusion is that as the number of labeled data decreases, the improvement of the framework is more obvious, which shows the reliability of the framework in the case of severely insufficient labeling.
[0144] Table 2 Semi - supervised multi - modal classification results
[0145]
[0146] Although some embodiments of the present disclosure have been shown and described, those skilled in the art should understand that these embodiments can be modified without departing from the principles and spirit of the present disclosure as defined by the claims and their equivalents.
Claims
1. A semi-supervised multi-modal classification method based on modality and policy complementarity, characterized in that, Including: A modality complementary learning module and a strategy complementary learning module; During the classification inference process, this method needs to simultaneously receive the images and texts in the input image-text pair data and hand them over to the encoders of their respective modalities. Among them: For the input image-text pair data, first, through the modality complementary learning module, the encoders of the text and the image are respectively applied, and an MLP model with two softmax functions is used to obtain the probability p, and single-modal predictions are made on the text or image data respectively; then, based on the extracted features or the single-modal prediction results, the results of the weighted score fusion method and the results of feature concatenation are obtained; And the results of the weighted score fusion and the feature concatenation results are input into the strategy complementary learning module to form a baseline model for semi-supervised multi-modal classification; The supervised loss L of the baseline model for semi-supervised multi-modal classification sup is expressed as: Among them, B x is the batch size of the sampled labeled data; At the output part, a label-level fusion method and a feature-level fusion method are used to obtain two different prediction results, and the average of the prediction probabilities of the two is taken to obtain the final class probability, and the class with the largest probability is selected as the classification result. Among them: For each unlabeled data u i Assign it a pseudo-label, for u i Prepare a weakly augmented version a(u i ) and a strongly augmented version A(u i ), and determine the pseudo-label of u by i and select the pseudo-label with high confidence to calculate the loss L u : Among them, B u is the batch size of unlabeled data, and τ is a fixed or variable threshold used to filter low-quality pseudo-labels; The pseudo-labeling method uses L sup + λL u or adds additional regularization as the optimized loss.
2. The semi-supervised multi-modal classification method based on modality and policy complementarity according to claim 1, wherein The specific method of the weighted score fusion is as follows: introduce the MLP model M w , and use the softmax function to obtain the reliability distribution w(z|t, v) of different modalities from the concatenated features [f(t), g(v)]. The specific method is as follows: Among them, is a parameter corresponding to the modality z'. For each image-text pair (t, v), a weighted fusion function is used: p S (y|t,v) = w(z = 0|t,v)p T (y|t) + w(z = 1|t,v)p V (y|v) This function is used for the score fusion branch.
3. A semi-supervised multi-modal classification method based on modal and policy complementarity according to claim 2, characterized in that To address the different training convergence speeds of strong and weak modalities, a label consistency guidance method is designed to calibrate the weight learning process using limited labels. Pseudo-labels and labeled data in the same batch are used to form a dataset D sub to train a modality reliability generator; The specific process is as follows: First, the prediction results are obtained through strong augmentation of the unlabeled input data: Among them, K is the set of categories, and k is a certain category in the category set; A(t) is the strongly augmented version of the text, A(v) is the strongly augmented version of the image, and p T p(k|A(t)) is the category prediction probability obtained from the strongly augmented version of the text, and p V p(k|A(v)) is the category prediction probability obtained from the strongly augmented version of the image; r t r is the predicted category result obtained from the text modality, and r v r is the category prediction result obtained from the image modality; Then, the supervision signal is generated according to the consistency of the pseudo-labels generated by the strongly augmented version and the weakly augmented version of the input: Among them, G(z|t,v) is the probability of the pseudo-label reliability distribution, is the pseudo-label, is the probability predicted to be the category based on the text input, is the probability predicted to be the category based on the image input; r t and r v are the category prediction results obtained from the above process; Correspondingly, there is G(z = 1|t, v) = 1 - G(z = 0|t, v); When r t or r v one of the modal predictions is consistent with the pseudo-label when they are consistent, the modal with consistent prediction is regarded as a reliable modal, and thus a probability score of 1 is assigned to this modal; When the prediction results r t and r v of both modalities are inconsistent with the pseudo labels, the prediction results of both modalities are more likely to be unreliable, and the contributions of both modalities are evenly divided; When the modal predictions r t and r v are both consistent with the pseudo-labels when they are both consistent with the pseudo-labels, the modal prediction with a higher confidence is regarded as a more reliable prediction, and the probability scores are calculated by performing a softmax operation based on their corresponding maximum prediction values; The KL divergence is used to make the generated reliability distribution w(z|t,v) close to the guiding value G(z|t,v), and this part of the loss L lcg is expressed as follows: where |D sub | is the size of a subset of the union set selected in the batch.
4. A semi-supervised multi-modal classification method based on modality and strategy complementarity according to claim 3, characterized in that To solve the problem that the reliability weight is small and the gradient generated by the unreliable modality is not sufficient to train the corresponding branch, a modality reliability guidance module is set in the modality complementary learning module, and the same data D as the label consistency guidance is adopted sub , the data subset is divided into two parts, and D t2i ∈D sub is introduced to represent the dataset with more reliable text modality, and D i2t ∈D sub is introduced to represent the dataset with more reliable image modality; for each data and label (t, v, y) in D sub , if r t is consistent with y, while r v is inconsistent with y, or both are consistent with y, but the maximum predicted value of p T is higher than that of p V , add these data to D t2i ; The KL divergence is used to complete the training process, and L mrg is introduced to represent the loss of the modality reliability guidance module: where |D t2i | and |D i2t | are the sizes of two sets.
5. A semi-supervised multi-modal classification method based on modal and policy complementarity according to claim 4, characterized in that The strategy complementary learning module introduces a consistency pseudo-label selection method to construct a dataset for guiding information; first, it aggregates the labeled data and unlabeled data with high confidence. For the unlabeled data, it retains the data for which the pseudo-labels predicted by the two strategies are consistent and the maximum probability of one of them is higher than the threshold η, denoted by U sub The subset of unlabeled data for a batch B is represented as: Among them, and are the pseudo-labels for weighted score fusion and feature concatenation; Then, use X sub ={(x i ,y i )|(x i ,y i ) ∈ B} represents all the labeled data of the same batch, indicating that the union set is D sub = U sub ∪ X sub ; Design a strategy complementary guidance module to apply the strategy complementarity of weighted score fusion and feature splicing, using D s2c ∈D sub to represent the data where feature splicing learns from score fusion, using D c2s ∈D sub to represent the opposite data, and use strong augmentation to collect and as the prediction results of unlabeled data for weighted score fusion and feature splicing. Here, K is the set of classes, k is a certain class in the class set; A(t) is the strongly augmented version of the text, A(v) is the strongly augmented version of the image; p S (k|(A(t),A(v)) is the predicted probability obtained by the score fusion method, p C (k|(A(t),A(v)) is the predicted probability obtained by the feature splicing method; r s is the predicted class result obtained by the score fusion method, r C is the predicted class result obtained by the feature splicing method; and use the non-augmented data to obtain r s and r C for the labeled data; The rules of D s2c and D c2s are expressed as: for each data and label (t, v, y) in D sub , if r s is consistent with y, but r C is inconsistent with y, or both are consistent with y, but the maximum predicted value of p S is higher than p C , add these data to D s2c to promote p C to learn from p S ; After that, the KL divergence is used to force the weaker strategy to learn from the stronger strategy, thereby obtaining L scg : where |D s2c | and |D c2s | are the sizes of two sets.