Multi-modal identification method and device based on noise tag, equipment, storage medium and program product
By distinguishing trusted and untrusted labels through the semantic trust boundary of noisy labels, and using the gradient of trusted labels to optimize the gradient of untrusted labels, the problem of low recognition accuracy of visual language models caused by noisy annotations in low-resource scenarios is solved, and the robustness of the model and the recognition accuracy are improved.
Patent Information
- Application Number
- CN202511134305.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-08-14
AI Technical Summary
Visual language models are affected by noisy annotations in low-resource scenarios, resulting in low recognition accuracy. Existing noise-resistant training methods ignore semantic structures and cannot effectively improve recognition accuracy.
The semantic trust boundary of the noise label is used to distinguish the trusted and untrusted labels, and the gradient of the trusted label is used to optimize the gradient projection of the untrusted label, thereby suppressing the optimization drift and optimizing the visual language model.
The recognition accuracy of the visual language model in low-resource scenarios is improved, and the model's robustness to noisy supervision is enhanced.
Smart Images

Figure CN120726421A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a noise label-based multimodal recognition method, apparatus, device, storage medium, and program product. Background Art
[0002] Currently, cue learning for Visual Language Models (VLMs) establishes a new paradigm for efficient transfer learning by separating parameter adaptation from feature extraction. Its potential in low-resource scenarios has been demonstrated in a variety of tasks, including pathology slide classification and hyperspectral object recognition.
[0003] In related technologies, just-in-time fine-tuning achieved by cue learning demonstrates superior performance in low-resource tasks, fundamentally due to the correctness of the supervisory signal. However, in practical applications, downstream datasets are often affected by noisy annotations caused by automated data collection or human error. This poses a significant challenge to just-in-time fine-tuning with sparse parameters, resulting in low recognition accuracy for visual language models trained with cue learning. Summary of the Invention
[0004] Based on this, it is necessary to provide a noise label-based multimodal recognition method, device, equipment, storage medium and program product that can improve recognition accuracy in response to the above technical problems.
[0005] In a first aspect, the present application provides a multimodal recognition method based on noise labels, comprising:
[0006] Obtain a noise sample dataset of a visual language model, wherein the noise sample dataset includes a plurality of image samples and sample labels corresponding to the plurality of image samples;
[0007] Dividing the sample labels into trusted labels and untrusted labels according to the semantic trust boundaries of the sample labels;
[0008] performing prompt learning on the visual language model according to the image samples with the trusted labels and the image samples with the untrusted labels, so as to optimize the visual language model;
[0009] Using the optimized visual language model, identify the image to be identified in the downstream task and determine the category of the image to be identified;
[0010] The visual language model is a model that aligns image and text representations in a shared embedding space through pre-training; the prompt learning is used to adjust the context vector in the visual language model through the gradient of the trusted label and the gradient of the untrusted label; the gradient of the untrusted label is generated after projection optimization based on the gradient of the trusted label; the context vector is used to generate text prompts for each category in the downstream task to map the text prompts to the text representation.
[0011] In one embodiment, before dividing the sample labels into trusted labels and untrusted labels according to the semantic trust boundaries of the sample labels, the method further includes:
[0012] Performing zero-shot prediction on the sample image of the sample label using the visual language model to obtain a zero-shot prediction result;
[0013] Determine a semantic trust boundary of the sample label based on the zero-sample prediction result.
[0014] In one embodiment, determining the semantic trust boundary of the sample label according to the zero-sample prediction result includes:
[0015] Determining a predicted category set corresponding to the sample label based on the zero-sample prediction result, the predicted category set including K categories with the highest probabilities obtained by the zero-sample prediction of the image sample corresponding to the sample label through the visual language model, where K is an integer greater than or equal to zero;
[0016] Determine a semantic trust boundary of the sample label based on the prediction category set corresponding to the sample label; wherein, the sample labels in the prediction category set are within the semantic trust boundary, and the sample labels not in the prediction category set are outside the semantic trust boundary.
[0017] In one embodiment, before determining the prediction category set corresponding to the sample label based on the zero-sample prediction result, the method further includes:
[0018] Determine the precision and recall corresponding to the set of predicted categories for different K values;
[0019] According to the precision and recall rates corresponding to the prediction category sets with different K values, the probability index scores corresponding to the prediction category sets with different K values are determined;
[0020] The K value corresponding to the highest probability index score is determined as the number of categories included in the predicted category set.
[0021] In one embodiment, optimizing the visual language model includes:
[0022] determining an average gradient of the trust labels;
[0023] Determining the cosine similarity corresponding to the gradient of the untrusted tag according to the average gradient of the trusted tag;
[0024] Optimizing the gradient of the untrusted label according to the cosine similarity corresponding to the gradient of the untrusted label;
[0025] The visual language model is optimized using the optimized gradient of the untrusted label and the gradient of the trusted label.
[0026] In one embodiment, optimizing the gradient of the untrusted label according to the cosine similarity corresponding to the gradient of the untrusted label includes:
[0027] If the cosine similarity is less than zero, projecting the gradient of the untrusted label into the trust alignment space corresponding to the average gradient of the trusted label;
[0028] The gradient of the untrusted label is optimized by removing the component of the gradient of the untrusted label in the direction of the average gradient in the trust alignment space.
[0029] In a second aspect, the present application further provides a multimodal recognition device based on noise labels, comprising:
[0030] An acquisition module is used to acquire a noise sample dataset of a visual language model, wherein the noise sample dataset includes a plurality of image samples and sample labels corresponding to the plurality of image samples;
[0031] a division module, configured to divide the sample labels into trusted labels and untrusted labels according to the semantic trust boundaries of the sample labels;
[0032] a learning module, configured to perform prompt learning on the visual language model based on the image samples with the trusted labels and the image samples with the untrusted labels, so as to optimize the visual language model;
[0033] A recognition module is used to use the optimized visual language model to identify the image to be recognized in the downstream task and determine the category of the image to be recognized;
[0034] The visual language model is a model that aligns image and text representations in a shared embedding space through pre-training; the prompt learning is used to adjust the context vector in the visual language model through the gradient of the trusted label and the gradient of the untrusted label; the gradient of the untrusted label is generated after projection optimization based on the gradient of the trusted label; the context vector is used to generate text prompts for each category in the downstream task to map the text prompts to the text representation.
[0035] In one embodiment, the segmentation module is further configured to perform zero-sample prediction on the sample image of the sample label using the visual language model to obtain a zero-sample prediction result; and determine the semantic trust boundary of the sample label based on the zero-sample prediction result.
[0036] In one embodiment, the partitioning module is further used to determine a prediction category set corresponding to the sample label based on the zero-sample prediction result, where the prediction category set includes the K categories with the highest probability obtained by the zero-sample prediction of the image sample corresponding to the sample label through the visual language model, where K is an integer greater than or equal to zero; and determine the semantic trust boundary of the sample label based on the prediction category set corresponding to the sample label; wherein the sample labels within the prediction category set are within the semantic trust boundary, and the sample labels not within the prediction category set are outside the semantic trust boundary.
[0037] In one embodiment, the partitioning module is further used to determine the precision and recall corresponding to the prediction category sets with different K values; determine the probability index scores corresponding to the prediction category sets with different K values based on the precision and recall corresponding to the prediction category sets with different K values; and determine the K value corresponding to the highest probability index score as the number of categories included in the prediction category set.
[0038] In one embodiment, the learning module is further used to determine the average gradient of the trusted label; determine the cosine similarity corresponding to the gradient of the untrusted label based on the average gradient of the trusted label; optimize the gradient of the untrusted label based on the cosine similarity corresponding to the gradient of the untrusted label; and optimize the visual language model using the optimized gradient of the untrusted label and the gradient of the trusted label.
[0039] In one embodiment, the learning module is further configured to project the gradient of the untrusted label into a trust alignment space corresponding to the average gradient of the trusted label if the cosine similarity is less than zero; and optimize the gradient of the untrusted label by removing the component of the gradient of the untrusted label in the direction of the average gradient in the trust alignment space.
[0040] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the multimodal recognition method based on noise labels of the first aspect is implemented.
[0041] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the multimodal recognition method based on noise labels of the first aspect.
[0042] In a fifth aspect, the present application further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the noise label-based multimodal recognition method of the first aspect.
[0043] The above-mentioned multimodal recognition method, device, equipment, storage medium and program product based on noise labels obtain a noise sample data set of a visual language model, which includes multiple image samples and sample labels corresponding to the multiple image samples; according to the semantic trust boundary of the sample labels, the sample labels are divided into trust labels and untrust labels; according to the image samples of the trust labels and the image samples of the untrust labels, the visual language model is prompted for learning respectively to optimize the visual language model; the optimized visual language model is used to identify the image to be identified in the downstream task and determine the category of the image to be identified; wherein the visual language model is a model that aligns the image and text representation in a shared embedding space through pre-training; prompt learning is used to adjust the context vector in the visual language model through the gradient of the trust label and the gradient of the untrust label; the gradient of the untrust label is generated after projection optimization based on the gradient of the trust label; the context vector is used to generate text prompts for each category in the downstream task to map the text prompts to the text representation. Since the semantic trust boundary of the sample label is used to distinguish the trusted label from the untrusted label, and then the gradient of the untrusted label is projected and optimized according to the gradient of the trusted label in the prompt learning, the optimization drift caused by the untrusted label is suppressed, thereby improving the accuracy of image category recognition using the visual language model. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.
[0045] Figure 1A diagram illustrating an application environment of a noise label-based multimodal recognition method provided in an embodiment of the present application;
[0046] Figure 2 A flowchart of a noise label-based multimodal recognition method provided in an embodiment of the present application;
[0047] Figure 3 A flowchart of another multimodal recognition method based on noise labels provided in an embodiment of the present application;
[0048] Figure 4 A flowchart of another noise label-based multimodal recognition method provided in an embodiment of the present application;
[0049] Figure 5 A structural block diagram of a noise label-based multimodal recognition device provided in an embodiment of the present application;
[0050] Figure 6 This is a diagram of the internal structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0051] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0052] First, the related technology will be described below.
[0053] Cued learning for visual language models establishes a new paradigm for efficient transfer learning by decoupling parameter adaptation from feature extraction. Specifically, due to the inherent properties of the dual-encoder architecture in visual language models—the frozen image encoder retains strong generalization capabilities, while cue learning semantically remaps downstream categories into a pre-trained semantic space—Context Optimization (CoOp) cue tuning can optimize only the textual cues of the visual language model. Consequently, the potential of cue learning for visual language models in low-resource scenarios has been demonstrated in a variety of tasks, including pathology slide classification and hyperspectral object recognition.
[0054] However, the superior performance of just-in-time fine-tuning enabled by hint learning on low-resource tasks stems from the correctness of the supervisory signal. In real-world applications, downstream datasets are often subject to noisy annotations, either from automated data collection or human error. This poses a significant challenge to just-in-time fine-tuning with sparse parameters. When asymmetric label noise is introduced, CoOp's performance degrades rapidly as the noise level increases.
[0055] It should be understood that in contrastive language-image pre-training (CLIP) of visual language models, the image and text encoders of the visual language model, even though they encode strong prior knowledge, are frozen. Therefore, textual prompts, as the only trainable component, are insufficient to distinguish semantic deviations or correct misalignments in the pre-training space in the presence of label noise. This leads to confusion in inter-class relationships and misleading optimization of the visual language model. Furthermore, the traditional cross-entropy loss function treats noisy and clean samples equally, further amplifying the negative impact of mislabeled data. Therefore, improving the robustness of fast tuning under noisy supervision has become a key bottleneck for its practical application.
[0056] In the related art, there are two methods for noise-resistant training of noisy labels. The first method directly filters out samples with unstable predictions through confidence modeling or uncertainty estimation. The second method uses robust loss functions such as generalized cross-entropy (GCE) to reduce noise interference by assigning lower gradient weights to low-confidence samples. However, these noise-resistant training methods in the related art ignore the semantic structure of the visual language model, especially the inter-class space constructed by CLIP.
[0057] The present application provides a multimodal recognition method, apparatus, device, storage medium and program product based on noise labels, which distinguishes trusted labels from untrusted labels through the semantic trust boundary of the noise labels, and then performs projection optimization on the gradient of the untrusted labels according to the gradient of the trusted labels in the prompt learning, thereby suppressing the optimization drift caused by the untrusted labels, and further improving the accuracy of image category recognition using the visual language model.
[0058] The following describes the application scenarios of the noise label-based multimodal recognition method involved in the embodiments of the present application.
[0059] The multimodal recognition method based on noise labels provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown, the terminal 102 communicates with the server 104 via a network. The data storage system can store data that the server 104 needs to process. The data storage system can be integrated on the server 104 or placed on the cloud or other network servers.
[0060] The server 104 may first obtain a noise sample dataset of the visual language model, which includes a plurality of image samples and sample labels corresponding to the plurality of image samples. Secondly, the server 104 divides the sample labels into trust labels and untrust labels according to the semantic trust boundary of the sample labels. Thirdly, the server 104 performs prompt learning on the visual language model based on the image samples with trust labels and the image samples with untrust labels to optimize the visual language model. Finally, the terminal 102 sends the image to be recognized in the downstream task to the server, and the server 104 uses the optimized visual language model to recognize the image to be recognized in the downstream task and determine the category of the image to be recognized.
[0061] Among them, the visual language model is a model that aligns image and text representations in a shared embedding space through pre-training; prompt learning is used to adjust the context vector in the visual language model through the gradient of the trust label and the gradient of the untrusted label; the gradient of the untrusted label is generated after projection optimization based on the gradient of the trust label; the context vector is used to generate text prompts for each category in the downstream task to map the text prompts to the text representation.
[0062] Terminal 102 may include, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices may include smart speakers, smart TVs, smart air conditioners, smart car devices, and projectors. Portable wearable devices may include smart watches, smart bracelets, and head-mounted devices. Head-mounted devices may include virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, and the like. Server 104 may be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server providing cloud computing services.
[0063] In an exemplary embodiment, Figure 2 As shown in the figure, a multimodal recognition method based on noise labels is provided. Figure 1 Taking the server in as an example, the multimodal recognition method based on noise labels includes S201-S204:
[0064] S201. Obtain a noise sample dataset for a visual language model.
[0065] It should be understood that the embodiments of this application are not limited to the aforementioned noise sample dataset. In some embodiments, the aforementioned noise sample dataset is a noise sample dataset corresponding to a downstream task, which is a specific application task in machine learning that needs to be solved based on a pre-trained model. Exemplary downstream tasks include case slide classification tasks and hyperspectral object recognition tasks.
[0066] In some embodiments, the noise sample dataset includes multiple image samples and sample labels corresponding to the multiple image samples. The sample labels are labels that may be wrong, inaccurate, or inconsistent with the true attributes of the samples, and are used to identify the category to which the corresponding sample images belong.
[0067] For example, if the sample dataset of the real label contains N samples, the sample dataset of the real label As shown in formula (1), when the real samples are changed into noise samples (i.e., image samples whose sample labels may be non-trusted labels), the noise sample dataset It can be shown as formula (2).
[0068] (1)
[0069] (2)
[0070] in, is the i-th sample image, N is the total number of samples, is the true label corresponding to the i-th sample image, , L is the total number of categories in the downstream task, is the noise label (i.e., non-trust label) corresponding to the i-th sample image.
[0071] It should be understood that a variety of reasons may cause label corruption, thereby causing the true label to become a noise label. For example, the true label to become a noise label can be expressed by formula (3).
[0072] (3)
[0073] in, is the probability that a sample with the true label i is incorrectly labeled as category j, is the true label, is the noise label.
[0074] In some embodiments, noise labels may include symmetric noise labels and asymmetric noise labels. Symmetric noise labels may be labels in which all incorrect categories are assigned the same probability of being noise. Asymmetric noise labels may be labels in which label flipping occurs only between specific category pairs, simulating real confusion between semantically or visually similar categories.
[0075] For example, the transfer matrix of the symmetric noise label can be shown as formula (4), and the transfer matrix of the asymmetric noise label can be shown as formula (5).
[0076] (4)
[0077] (5)
[0078] in, is the noise rate. When , the label is clean; when , the labels are completely randomized. represents classes that are visually similar to class i.
[0079] It should be understood that the noise rate is used to characterize the degree or proportion of noise present in the noise sample dataset. The noise labels in the noise sample dataset can be symmetric noise labels or asymmetric noise labels. The noise level can be set according to the actual situation, for example, 0.125, 0.25, 0.75, etc.
[0080] S202: Divide the sample labels into trust labels and untrust labels according to the semantic trust boundaries of the sample labels.
[0081] In this step, after the server obtains the noise sample dataset of the visual language model, it can divide the sample labels into trusted labels and untrusted labels according to the semantic trust boundaries of the sample labels.
[0082] The semantic trust boundary of a sample label is a boundary constructed by the degree of proximity between the sample label and the zero-sample prediction. Sample labels within the semantic trust boundary can be considered trust labels, and sample labels outside the semantic trust boundary can be considered untrust labels.
[0083] Among them, the trusted label can be trusted as the true label; the untrusted label cannot be trusted as the true label, that is, the untrusted label can be considered as the above-mentioned noise label.
[0084] In some embodiments, before classifying sample labels into trusted labels and untrusted labels based on their semantic trust boundaries, the server may first use a visual language model to perform zero-shot prediction on sample images of the sample labels to obtain zero-shot prediction results. Subsequently, the server determines the semantic trust boundaries of the sample labels based on the zero-shot prediction results.
[0085] In particular, zero-shot prediction refers to the prediction or classification of the category of the downstream task by the visual language model only through prior information without training for the downstream task.
[0086] It should be understood that in zero-shot prediction, even if the top-1 prediction is incorrect, the true label often appears in the top-ranked predictions. Therefore, the semantic trust boundary can be constructed by the set of predicted categories of the K category groups with the highest probability obtained by zero-shot prediction.
[0087] In some embodiments, the server may first determine a set of predicted categories corresponding to the sample label based on the zero-sample prediction result, and then determine a semantic trust boundary of the sample label based on the set of predicted categories corresponding to the sample label.
[0088] The predicted category set includes the K categories with the highest probabilities obtained by zero-shot prediction of the image sample corresponding to the sample label using the visual language model, where K is an integer greater than or equal to zero. Exemplarily, the predicted category set may be a Top-K category set, which includes the K categories with the highest probabilities.
[0089] Correspondingly, if the sample label in the predicted category set is within the semantic trust boundary, it means that the sample label is trusted, and if the sample label not in the predicted category set is outside the semantic trust boundary, it means that the sample label is not trusted.
[0090] For example, if the image sample is , the probability distribution obtained by zero-sample prediction is , correspondingly, the Top-K category set It can be shown as formula (6).
[0091] (6)
[0092] For example, before training, we can use the static template "photos of [CLASS]" as a text prompt for Top-K evaluation to establish an initial semantic consistency reference. For example, after training begins, we switch to text prompts generated by learnable context vectors. The text prompts generated by the learnable context vectors are continuously updated during the visual language model training process, as shown in Formula (7).
[0093] (7)
[0094] Among them, m∈[1,M], It can be the mth learnable context vector, which is continuously updated during the training process. A vector of category names for downstream tasks.
[0095] For example, as the representation ability of the visual language model improves during training, the accuracy of semantic consistency estimation will also improve. If the sample label of the image sample is included in the Top-K category set If the image sample’s label is not included in the Top-K category set, then the image sample is considered to be semantically consistent with the model prediction and is classified as a trust label. If the value is less than 0, the image sample is considered to be semantically inconsistent with the model prediction and can be used as an untrusted label.
[0096] K is a hyperparameter that controls the number of candidate categories. The hyperparameter K corresponding to the Top-K category set effectively acts as a classification threshold in semantic verification. Therefore, the choice of K directly affects the distinction between trusted and untrusted labels. This hyperparameter K can be determined by the highest probability metric F1 score.
[0097] In some embodiments, the server may first determine the precision and recall corresponding to the prediction category sets for different values of K. Subsequently, the server determines the probability index scores corresponding to the prediction category sets for different values of K based on the precision and recall corresponding to the prediction category sets for different values of K. Finally, the server determines the K value corresponding to the highest probability index score as the number of categories included in the prediction category set.
[0098] Example rows, the accuracy of the predicted category set corresponding to different K values The recall rate corresponding to the prediction category set with different K values can be determined by formula (8): It can be determined by formula (9). After determining the precision and recall, the probability index score corresponding to the prediction category set with different K values can be determined by formula (10). Then, the optimal K value corresponding to the highest probability index score can be selected by formula (11) to achieve a balanced trade-off, that is, the number of categories included in the prediction category set.
[0099] (8)
[0100] (9)
[0101] (10)
[0102] (11)
[0103] Among them, K is a hyperparameter that controls the number of candidate categories and is used to characterize the number of categories contained in the predicted category set; For accuracy, is the recall rate, Score the probability indicator, is a non-trusted label, is the sample label, is a noise sample dataset.
[0104] In some embodiments, after determining K, a soft label may be assigned to the sample label. For example, the soft label may be assigned to the sample label according to formula (12).
[0105] (12)
[0106] in, is the soft label, L is the total number of categories of downstream tasks, is the smoothing factor, , K is a hyperparameter that controls the number of candidate categories and is used to characterize the number of categories contained in the predicted category set.
[0107] For example, by assigning uniform Probability, assign zero probability to all other classes to initialize the soft labels. At this point, a smoothing factor can be further applied to obtain the final label distribution. The probability mass of the Top-K categories is , the probability mass of the remaining categories is In this way, over-trust in potentially noisy targets can be prevented to enhance robustness.
[0108] The following describes the false positive rate of label division under the semantic trust boundary of sample labels.
[0109] It should be understood that the above-mentioned false alarm rate may be the probability of incorrectly classifying a non-trusted label as a trusted label under the semantic trust boundary.
[0110] For example, a label distrust setting is defined by the transformation matrix Control, where entries , total distrust rate Correspondingly, the bound of the false alarm rate can be expressed as formula (13).
[0111] (13)
[0112] S203 : Perform prompt learning on the visual language model based on the image samples with the trusted labels and the image samples with the untrusted labels, so as to optimize the visual language model.
[0113] In this step, after the server divides the sample labels into trusted labels and untrusted labels according to the semantic trust boundaries of the sample labels, the visual language model can be prompted for learning based on the image samples of the trusted labels and the image samples of the untrusted labels to optimize the visual language model.
[0114] Among them, the visual language model is a model that aligns image and text representations in a shared embedding space through pre-training; prompt learning is used to adjust the context vector in the visual language model through the gradient of the trust label and the gradient of the untrust label; the context vector is used to generate text prompts for each category in the downstream task to map the text prompts to the text representation.
[0115] For example, the visual language model aligns image and text representations in a shared embedding space by performing large-scale comparative learning on paired image-text data during pre-training. In the downstream classification task, prompt learning introduces a set of learnable context vectors to construct text prompts describing each category in the downstream task, thereby optimizing the visual language model. The context vector can be shown as formula (7). The sample image x is passed through the image encoder in the visual language model. Processing to obtain visual features ; Text prompt Text Encoder in Vision-Language Model Processing to obtain corresponding text features The cosine similarity between these two features is used to measure the semantic alignment between the sample image and the text representation. Correspondingly, the probability of classifying the sample image x into category i through the visual language model is As shown in formula (14).
[0116] (14)
[0117] in, is the probability that the sample image x is classified into category i, is the visual feature of the sample image, Text prompt The text features of L are the total number of categories of the downstream task, and T is the temperature parameter, which is used to control the smoothness of the output. In the L-way classification setting, the visual language model generates the final prediction vector .
[0118] It should be understood that during the optimization process, hint learning only optimizes the context vector while keeping the image encoder and text encoder in a frozen state.
[0119] In some embodiments, the gradient of the untrusted label is generated after performing projection optimization based on the gradient of the trusted label.
[0120] In an embodiment of the present application, through the semantic trust boundary, trust labels and untrust labels can be effectively identified from sample labels. Trust labels can be used for reliable supervised learning, but the processing of untrust labels remains a key challenge. Simply discarding sample images corresponding to untrusted labels or uniformly reducing their influence may result in information loss or unsatisfactory optimization results. Therefore, when performing prompt learning, it is necessary to retain potentially useful learning signals from sample images corresponding to untrusted labels while preventing the introduction of harmful gradient directions.
[0121] Accordingly, in an embodiment of the present application, trust-aligned gradient projection is introduced. Through the gradient-level denoising mechanism of trust-aligned gradient projection, the semantic supervision number of the sample corresponding to the trust label is used to construct a trust alignment space in the average gradient direction of the trust label, and the gradient of the non-trust label is projected into the trust alignment space, thereby optimizing the gradient of the non-trust label.
[0122] In some embodiments, the server may first determine the average gradient of the trusted tags. Next, the server may determine the cosine similarity of the gradients of the untrusted tags based on the average gradient of the trusted tags. Third, the server may optimize the gradients of the untrusted tags based on the cosine similarity of the gradients of the untrusted tags. Finally, the server may use the optimized gradients of the untrusted tags and the gradients of the trusted tags to optimize the visual language model.
[0123] For example, the image sample set corresponding to the trust label can be , the image sample set corresponding to the non-trust label can be , the average gradient of the trust label can be shown as formula (15). For each untrusted sample, its gradient can be calculated according to the corresponding soft label, as shown in formula (16). Subsequently, the cosine similarity corresponding to the gradient of the untrusted label is calculated by formula (17).
[0124] (15)
[0125] (16)
[0126] (17)
[0127] in, is the training loss, are model parameters, is the average gradient of the trust label. is the gradient of the i-th non-trusted label, is the cosine similarity of the i-th non-trusted label.
[0128] In some embodiments, the gradient of an untrusted tag can be conditionally adjusted based on the gradient consistency principle. If the cosine similarity is less than zero, the server projects the gradient of the untrusted tag into the trust alignment space corresponding to the average gradient of the trusted tag. The server then optimizes the gradient of the untrusted tag by removing the component of the untrusted tag's gradient in the direction of the average gradient in the trust alignment space. If the cosine similarity is greater than or equal to zero, the server can directly use the gradient of the untrusted tag as the optimized gradient of the untrusted tag.
[0129] For example, if , then retain the gradient of the untrusted label as the gradient of the optimized untrusted label; if , then the gradient of the non-trusted label is projected into the trust alignment space corresponding to the average gradient of the trusted label, suppressing or removing the gradient in The component in the direction of the untrusted label is reduced to mitigate its adverse effect on model optimization. For example, the gradient of the optimized untrusted label can be shown as formula (18). For example, the final optimization direction of the visual language model is determined by the optimized gradient of the untrusted label and the gradient of the trusted label, which can be shown as formula (19).
[0130] (18)
[0131] (19)
[0132] in, is the gradient clipping coefficient, which is used to control the projection-based suppression strength. For example, The value of can be 1. is the gradient of the optimized untrusted label.
[0133] In this application, a gradient consistency constraint is introduced in the optimization process of the visual language model to suppress harmful updates from untrusted labels and mitigate convergence drift by aligning the gradient direction with the supervision of trusted labels.
[0134] In some embodiments, the loss function of the visual language model may be a noise-aware supervised loss function, which may enhance the training robustness to untrusted samples.
[0135] For example, the loss function of noise-aware supervision can be shown as formula (20). When q = 0, the formula is simplified to the standard cross entropy, and when q = 1, the formula is linearly equivalent to the mean absolute error. For example, q can be set to 0.5 to strike a balance between robustness and optimization efficiency, so that the loss function of noise-aware supervision provides a stable pseudo-supervision signal, which helps to mitigate the adverse effects of untrusted labels during training.
[0136] (20)
[0137] in, is the loss function of noise-aware supervision, q is the noise parameter, .
[0138] S204: Use the optimized visual language model to identify the image to be identified in the downstream task and determine the category of the image to be identified.
[0139] In this step, after the optimized visual language model is obtained, the optimized visual language model can be used to identify the image to be identified in the downstream task and determine the category of the image to be identified.
[0140] In some embodiments, the terminal can upload the image to be recognized in the downstream task to the server. After receiving the image to be recognized, the server can input the image to be recognized into the optimized visual language model and obtain the category recognition result output by the optimized visual language model.
[0141] Exemplarily, downstream tasks include case slide classification tasks, hyperspectral object recognition tasks, etc., which are not limited in the embodiments of the present application.
[0142] The embodiment of the present application provides a multimodal recognition method based on noise labels, which obtains a noise sample data set of a visual language model, wherein the noise sample data set includes multiple image samples and noise labels corresponding to the multiple image samples; the noise labels are divided into trust labels and untrust labels according to the semantic trust boundaries of the noise labels; the visual language model is respectively prompted with learning based on the image samples with trust labels and the image samples with untrust labels to optimize the visual language model; the optimized visual language model is used to identify the image to be identified in the downstream task and determine the category of the image to be identified; wherein the visual language model is a model that aligns image and text representations in a shared embedding space through pre-training; the prompt learning is used to adjust the context vector in the visual language model through the gradient of the trust label and the gradient of the untrust label; the gradient of the untrust label is generated after projection optimization based on the gradient of the trust label; the context vector is used to generate text prompts for each category in the downstream task to map the text prompts to the text representation. Because the semantic trust boundary of the noise label is used to distinguish the trusted label from the untrusted label, and then the gradient of the untrusted label is projected and optimized according to the gradient of the trusted label in the prompt learning, the optimization drift caused by the untrusted label is suppressed, thereby improving the accuracy of image category recognition using the visual language model.
[0143] The following describes how to determine the semantic trust boundary of sample labels. Figure 3 A flow chart of another multimodal recognition method based on noise labels provided in an embodiment of the present application is shown as follows: Figure 3 As shown, the noise label-based multimodal recognition method includes S301-S307:
[0144] S301: Obtain a noise sample dataset of a visual language model, where the noise sample dataset includes a plurality of image samples and sample labels corresponding to the plurality of image samples.
[0145] S302: Use the visual language model to perform zero-shot prediction on the sample image of the sample label to obtain a zero-shot prediction result.
[0146] S303. Determine a prediction category set corresponding to the sample label based on the zero-sample prediction result. The prediction category set includes K categories with the highest probabilities obtained by the zero-sample prediction of the image sample corresponding to the sample label using the visual language model, where K is an integer greater than or equal to zero.
[0147] S304: Determine the semantic trust boundary of the sample label based on the prediction category set corresponding to the sample label, wherein the sample label in the prediction category set is within the semantic trust boundary, and the sample label not in the prediction category set is outside the semantic trust boundary.
[0148] S305 : Divide the sample labels into trust labels and untrust labels according to the semantic trust boundaries of the sample labels.
[0149] S306 : Perform prompt learning on the visual language model based on the image samples with the trusted labels and the image samples with the untrusted labels, so as to optimize the visual language model.
[0150] Among them, the visual language model is a model that aligns image and text representations in a shared embedding space through pre-training; prompt learning is used to adjust the context vector in the visual language model through the gradient of the trust label and the gradient of the untrusted label; the gradient of the untrusted label is generated after projection optimization based on the gradient of the trust label; the context vector is used to generate text prompts for each category in the downstream task to map the text prompts to the text representation.
[0151] S307: Use the optimized visual language model to identify the image to be identified in the downstream task and determine the category of the image to be identified.
[0152] In an embodiment of the present application, a semantic trust boundary is introduced to distinguish reliable label areas and divide trusted labels from untrusted labels. By explicitly integrating the semantic prior of the visual language model, the model's sensitivity to label noise is enhanced, thereby improving the accuracy of the division between trusted labels and untrusted labels, and further improving the recognition accuracy.
[0153] The following describes how to perform projection optimization on the gradient of the untrusted label based on the gradient of the trusted label. Figure 4 A flow chart of another multimodal recognition method based on noise labels provided in an embodiment of the present application is shown as follows: Figure 4 As shown, the noise label-based multimodal recognition method includes S401-S408:
[0154] S401: Obtain a noise sample dataset of a visual language model, where the noise sample dataset includes a plurality of image samples and sample labels corresponding to the plurality of image samples.
[0155] S402: Divide the sample labels into trust labels and untrust labels according to the semantic trust boundaries of the sample labels.
[0156] S403 : Perform prompt learning on the visual language model based on the image samples with the trusted label and the image samples with the untrusted label.
[0157] Among them, the visual language model is a model that aligns image and text representations in a shared embedding space through pre-training; prompt learning is used to adjust the context vector in the visual language model through the gradient of the trust label and the gradient of the untrusted label; the gradient of the untrusted label is generated after projection optimization based on the gradient of the trust label; the context vector is used to generate text prompts for each category in the downstream task to map the text prompts to the text representation.
[0158] S404: Determine the average gradient of the trust label.
[0159] S405 : Determine the cosine similarity corresponding to the gradient of the untrusted tag according to the average gradient of the trusted tag.
[0160] S406 : Optimize the gradient of the untrusted label according to the cosine similarity corresponding to the gradient of the untrusted label.
[0161] S407: Optimize the visual language model using the optimized gradient of the untrusted label and the gradient of the trusted label.
[0162] S408: Use the optimized visual language model to identify the image to be identified in the downstream task and determine the category of the image to be identified.
[0163] The noisy label-based multimodal recognition method provided in the embodiments of the present application generally achieves the best performance in all evaluation benchmarks and noise configurations. Under severe noise conditions (especially when the label damage rate is as high as 75%), the above-mentioned noisy label-based multimodal recognition method significantly outperforms all competing methods. For example, under 75% asymmetric noise conditions, the above-mentioned noisy label-based multimodal recognition method performs 10.8% higher than the Natural Language Prompt (NLPrompt) method on the noisy sample dataset (Flowers102), and 13.21% higher than the NLPromptt method on the noisy sample dataset (DTD). In sharp contrast, the performance of the methods in the related art (for example, the NLPromptt method) drops sharply as the noise rate increases, and the accuracy rate usually drops below 20%, which highlights its strong resistance to label noise accumulation.
[0164] In an embodiment of the present application, a trusted supervision space is constructed by aggregating the average gradient direction of the trust label, that is, the trust alignment space corresponding to the average gradient of the trust label. Then, the gradient of the non-trusted label is projected into the trust alignment space corresponding to the average gradient of the trust label, and only the gradient component of the non-trusted label gradient that is semantically aligned with the average gradient of the trust label is retained. The visual language model trained in this way will not discard or reduce the weight of the sample image corresponding to the non-trusted label, thereby effectively suppressing semantic interference while maintaining effective supervision, thereby significantly enhancing the robustness of fast tuning under label noise conditions, and thereby improving the accuracy of multimodal recognition based on noise labels.
[0165] The embodiment of the present application provides a multimodal recognition method based on noise labels, which obtains a noise sample data set of a visual language model, wherein the noise sample data set includes multiple image samples and sample labels corresponding to the multiple image samples; the sample labels are divided into trust labels and untrust labels according to the semantic trust boundaries of the sample labels; the visual language model is prompted with learning based on the image samples with trust labels and the image samples with untrust labels, respectively, to optimize the visual language model; the optimized visual language model is used to identify images to be identified in downstream tasks and determine the category of the images to be identified; wherein the visual language model is a model that aligns image and text representations in a shared embedding space through pre-training; the prompt learning is used to adjust the context vector in the visual language model through the gradient of the trust label and the gradient of the untrust label; the gradient of the untrust label is generated after projection optimization based on the gradient of the trust label; the context vector is used to generate text prompts for each category in the downstream task to map the text prompts to the text representation. Since the semantic trust boundary of the sample label is used to distinguish the trusted label from the untrusted label, and then the gradient of the untrusted label is projected and optimized according to the gradient of the trusted label in the prompt learning, the optimization drift caused by the untrusted label is suppressed, thereby improving the accuracy of image category recognition using the visual language model.
[0166] It should be understood that, although the steps in the flowcharts of the above embodiments are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the flowcharts of the above embodiments may include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these steps or stages is not necessarily to be performed in sequence, but can be performed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0167] Based on the same inventive concept, an embodiment of the present application further provides a noise tag-based multimodal recognition device for implementing the noise tag-based multimodal recognition method mentioned above. The implementation solution provided by this device is similar to the implementation solution described in the above method. Therefore, the specific limitations of one or more noise tag-based multimodal recognition device embodiments provided below can be found in the limitations of the noise tag-based multimodal recognition method above and will not be repeated here.
[0168] In an exemplary embodiment, Figure 5 As shown, a multimodal recognition device 500 based on noise labels is provided, comprising: an acquisition module 501, a division module 502, a learning module 503 and a recognition module 504, wherein:
[0169] The acquisition module 501 is used to acquire a noise sample dataset of a visual language model, where the noise sample dataset includes a plurality of image samples and sample labels corresponding to the plurality of image samples.
[0170] The division module 502 is configured to divide the sample labels into trusted labels and untrusted labels according to the semantic trust boundaries of the sample labels.
[0171] The learning module 503 is configured to perform prompt learning on the visual language model based on the image samples with the trusted labels and the image samples with the untrusted labels, so as to optimize the visual language model.
[0172] The recognition module 504 is configured to use the optimized visual language model to recognize the image to be recognized in the downstream task and determine the category of the image to be recognized.
[0173] Among them, the visual language model is a model that aligns image and text representations in a shared embedding space through pre-training; prompt learning is used to adjust the context vector in the visual language model through the gradient of the trust label and the gradient of the untrusted label; the gradient of the untrusted label is generated after projection optimization based on the gradient of the trust label; the context vector is used to generate text prompts for each category in the downstream task to map the text prompts to the text representation.
[0174] In one embodiment, the segmentation module 502 is further configured to perform zero-sample prediction on the sample image of the sample label using the visual language model to obtain a zero-sample prediction result; and determine the semantic trust boundary of the sample label based on the zero-sample prediction result.
[0175] In one embodiment, the partitioning module 502 is further used to determine a prediction category set corresponding to the sample label based on the zero-sample prediction result, where the prediction category set includes the K categories with the highest probability obtained by the zero-sample prediction of the image sample corresponding to the sample label through the visual language model, where K is an integer greater than or equal to zero; and determine the semantic trust boundary of the sample label based on the prediction category set corresponding to the sample label; wherein the sample label within the prediction category set is within the semantic trust boundary, and the sample label not within the prediction category set is outside the semantic trust boundary.
[0176] In one embodiment, the partitioning module 502 is further used to determine the precision and recall corresponding to the prediction category sets with different K values; determine the probability index scores corresponding to the prediction category sets with different K values based on the precision and recall corresponding to the prediction category sets with different K values; and determine the K value corresponding to the highest probability index score as the number of categories included in the prediction category set.
[0177] In one embodiment, the learning module 503 is further configured to determine an average gradient of a trusted tag; determine a cosine similarity corresponding to a gradient of an untrusted tag based on the average gradient of the trusted tag; optimize the gradient of the untrusted tag based on the cosine similarity corresponding to the gradient of the untrusted tag; and optimize the visual language model using the optimized gradient of the untrusted tag and the gradient of the trusted tag.
[0178] In one embodiment, the learning module 503 is further configured to project the gradient of the untrusted label into the trust alignment space corresponding to the average gradient of the trusted label if the cosine similarity is less than zero; and optimize the gradient of the untrusted label by removing the component of the gradient of the untrusted label in the direction of the average gradient in the trust alignment space.
[0179] Each module in the noise label-based multimodal recognition device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a memory in the computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0180] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 6As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store XX data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a multimodal recognition method based on noise labels is implemented.
[0181] Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0182] In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the above-mentioned noise label-based multimodal recognition method when executing the computer program.
[0183] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the multimodal recognition method based on noise labels is implemented.
[0184] In one embodiment, a computer program product is provided, including a computer program, which implements the above-mentioned noise label-based multimodal recognition method when executed by a processor.
[0185] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), quantum computing-based data processing logic devices, artificial intelligence (AI) processors, and the like.
[0186] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0187] The above embodiments merely illustrate several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art may make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A multimodal recognition method based on noise labels, characterized in that: The method comprises: Obtain a noise sample dataset of a visual language model, wherein the noise sample dataset includes a plurality of image samples and sample labels corresponding to the plurality of image samples; Dividing the sample labels into trusted labels and untrusted labels according to the semantic trust boundaries of the sample labels; performing prompt learning on the visual language model according to the image samples with the trusted labels and the image samples with the untrusted labels, so as to optimize the visual language model; Using the optimized visual language model, identify the image to be identified in the downstream task and determine the category of the image to be identified; The visual language model is a model that aligns image and text representations in a shared embedding space through pre-training; the prompt learning is used to adjust the context vector in the visual language model through the gradient of the trusted label and the gradient of the untrusted label; the gradient of the untrusted label is generated after projection optimization based on the gradient of the trusted label; the context vector is used to generate text prompts for each category in the downstream task to map the text prompts to the text representation.
2. The method according to claim 1, characterized in that Before dividing the sample labels into trusted labels and untrusted labels according to the semantic trust boundaries of the sample labels, the method further includes: Performing zero-shot prediction on the sample image of the sample label using the visual language model to obtain a zero-shot prediction result; Determine a semantic trust boundary of the sample label based on the zero-sample prediction result.
3. The method according to claim 2, characterized in that Determining the semantic trust boundary of the sample label according to the zero-sample prediction result includes: Determining a predicted category set corresponding to the sample label based on the zero-sample prediction result, the predicted category set including K categories with the highest probabilities obtained by the zero-sample prediction of the image sample corresponding to the sample label through the visual language model, where K is an integer greater than or equal to zero; Determine a semantic trust boundary of the sample label based on the prediction category set corresponding to the sample label; wherein, the sample labels in the prediction category set are within the semantic trust boundary, and the sample labels not in the prediction category set are outside the semantic trust boundary.
4. The method according to claim 3, characterized in that Before determining the prediction category set corresponding to the sample label according to the zero-sample prediction result, the method further includes: Determine the precision and recall corresponding to the set of predicted categories for different K values; According to the precision and recall rates corresponding to the prediction category sets with different K values, the probability index scores corresponding to the prediction category sets with different K values are determined; The K value corresponding to the highest probability index score is determined as the number of categories included in the predicted category set.
5. The method according to claim 1, wherein Optimizing the visual language model includes: determining an average gradient of the trust labels; Determining the cosine similarity corresponding to the gradient of the untrusted tag according to the average gradient of the trusted tag; Optimizing the gradient of the untrusted label according to the cosine similarity corresponding to the gradient of the untrusted label; The visual language model is optimized using the optimized gradient of the untrusted label and the gradient of the trusted label.
6. The method according to claim 5, characterized in that Optimizing the gradient of the untrusted label according to the cosine similarity corresponding to the gradient of the untrusted label includes: If the cosine similarity is less than zero, projecting the gradient of the untrusted label into the trust alignment space corresponding to the average gradient of the trusted label; The gradient of the untrusted label is optimized by removing the component of the gradient of the untrusted label in the direction of the average gradient in the trust alignment space.
7. A multimodal recognition device based on noise labels, characterized in that: The device comprises: An acquisition module is used to acquire a noise sample dataset of a visual language model, wherein the noise sample dataset includes a plurality of image samples and sample labels corresponding to the plurality of image samples; a division module, configured to divide the sample labels into trusted labels and untrusted labels according to the semantic trust boundaries of the sample labels; a learning module, configured to perform prompt learning on the visual language model based on the image samples with the trusted labels and the image samples with the untrusted labels, so as to optimize the visual language model; A recognition module is used to use the optimized visual language model to identify the image to be recognized in the downstream task and determine the category of the image to be recognized; The visual language model is a model that aligns image and text representations in a shared embedding space through pre-training; the prompt learning is used to adjust the context vector in the visual language model through the gradient of the trusted label and the gradient of the untrusted label; the gradient of the untrusted label is generated after projection optimization based on the gradient of the trusted label; the context vector is used to generate text prompts for each category in the downstream task to map the text prompts to the text representation.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
SQL injection attack protection method and device based on zero-trust framework
CN119728167A
Cited By
Image processing method and device, computer equipment and storage medium
CN121527596A